EventsThe 9th International Electronic Conference on Water Sciences
Published
This submission belongs to the session S7. Remote Sensing, Artificial Intelligence and New Technologies in Water Sciences of the event The 9th International Electronic Conference on Water Sciences
Published date
06 Nov, 2025
Academic Editor
author-avatarNikiforos Samarinas
Citation
Dimitrios Koulouris, Nikolaos Malamos, Soil Moisture Time Series Gap-Filling Using Random Forest Machine Learning Models: A Case Study in the Arta Plain, in Proceedings of The 9th International Electronic Conference on Water Sciences, 11 November–14 November 2025, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Soil Moisture Time Series Gap-Filling Using Random Forest Machine Learning Models: A Case Study in the Arta Plain

image
image
1. Department of Agriculture; University of Patras; Messolonghi; 30200; Greece, Greece
2. Department of Natural Resources Development & Agricultural Engineering; Agricultural University of Athens; Athens; 11855 ,Greece, Greece
Abstract

Soil moisture content (SMC) is a key environmental variable which influences numerous hydrological and ecological processes. However, the complex and dynamic nature of SMC makes it difficult to estimate. Moreover, invalid SMC measurements and data gaps in sensor-based SMC monitoring are common occurrences due to various reasons. This study investigates the effectiveness of the Random Forest (RF) machine learning algorithm in reconstructing missing SMC time series at depths of 10 cm, 30 cm, and 50 cm at two agricultural sites in the Arta plain, Greece. Input data included existing SMC time series at alternative depths and the NDVI vegetation index derived from Sentinel-2 satellite data. RF models were trained using daily SMC data from 2020 to 2021 and validated with 2022 observations. Model performance was evaluated using the Nash–Sutcliffe efficiency (NSE) and Root Mean Square Error (RMSE). The results demonstrated high predictive accuracy, with NSE values up to 0.98 and RMSE as low as 0.33 m³/m³. The best results were achieved when two SMC series were used as inputs. NDVI contributed less to model improvement, possibly because the NDVI daily time series is derived through temporal interpolation, as the NDVI values are not originally available on a daily basis. In addition, some NDVI values are discarded when the satellite image has more than 10% cloud cover. Overall, the study confirms that RF models are effective for imputing missing SMC data and can support irrigation management by reconstructing reliable soil moisture records even with limited sensor information.

Keywords
Soil Moisture Content
Random Forest
Machine Learning
Vegetation Index
NDVI
Oral Presentation
Spatio-Temporal Drought Assessment in the Pinios River Basin Using Ground Observations and Satellite Data.
Land Use and Land Cover Analysis and Prediction Using Machine Learning Approach: A Case Study of Gaibandha District, Bangladesh