Pesticide contamination of surface waters represents a persistent environmental challenge, particularly in agriculturally intensive regions where runoff and spray drift continuously introduce compounds into freshwater systems. Despite growing regulatory attention, routine monitoring programs remain limited in both spatial coverage and the number of substances analyzed, resulting in significant blind spots regarding actual contamination levels. This study explored whether Random Forest machine learning models could help bridge these gaps using monitoring data from Bretagne, France, a region where roughly 65% of land is under agricultural use.
Two models were developed, one predicting pesticide occurrence and one predicting concentrations, trained on predictors including land use, seasonality, chemical identity, and pesticide class. Both models showed strong predictive performance, with the concentration model accounting for over 99% of variance in independent test data. Compound-specific variables, particularly chemical identity and pesticide class, proved far more informative than spatial predictors like land use, though seasonal and landscape factors still contributed meaningfully. Applying the models to an expanded dataset covering unmonitored site-pesticide combinations suggested that actual contamination levels may be nearly three times higher than what conventional monitoring programs currently detect.
These results point to a substantial underestimation of pesticide presence in freshwater and support the use of data-driven approaches as a practical complement to field monitoring in environmental risk assessment.