EventsThe 1st International Online Conference on Earth Science
Published
This submission belongs to the session S7. Air Quality and Climate Pollutants of the event The 1st International Online Conference on Earth Science
Published date
31 Aug, 2026
Academic Editor
author-avatarAlexander A. Baklanov
Citation
Moumita Saha, Piyush Goenka, Sorbojit Mondal, Ricky Dey, Yuvraj Prasad, Estimating Ground-Level PM2.5 from Open Satellite and Meteorological Data Using Explainable Machine Learning, in Proceedings of The 1st International Online Conference on Earth Science, 2 September–4 September 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Estimating Ground-Level PM2.5 from Open Satellite and Meteorological Data Using Explainable Machine Learning

Sorbojit Mondal 2
Ricky Dey 2
Yuvraj Prasad 2
1. Sustainability, Tata Consultancy Services (TCS), Kolkata, India
2. Techno International New Town, Kolkata, India
Abstract

Introduction: Fine particulate matter (PM2.5) is a major climate-relevant air pollutant linked to public-health burden, ecosystem stress, and regional haze. Ground monitoring networks remain sparse and unevenly distributed across many urban areas, limiting continuous surveillance required for mitigation planning. Satellite-derived aerosol optical depth (AOD) provides spatially extensive coverage, but the AOD–PM2.5 relationship is nonlinear
and strongly modulated by meteorology and land cover. Machine learning models that fuse open-access satellite, meteorological, and ground observations can improve gap-filled PM2.5 estimation, yet reproducible comparative frameworks built entirely on public archives are still limited.

Methods: This study develops an explainable machine learning workflow to estimate daily ground-level PM2.5 using open-access earth-observation and air-quality data. Ground concentrations are compiled from OpenAQ and national monitoring records; satellite predictors include MODIS/MERRA-2 AOD; meteorological covariates (temperature, relative humidity, wind speed, precipitation) are obtained from ERA5 reanalysis; vegetation context
is represented using NDVI. After temporal alignment and quality control, linear, Support Vector Regression, Random Forest, XGBoost, and LightGBM models are trained and evaluated using chronological and spatial cross-validation (RMSE, MAE, MAPE, R2). SHAP analysis quantifies feature contributions and supports interpretation of pollution–weather interactions.

Results: Ensemble tree models are expected to outperform linear AOD–PM2.5 baselines by capturing nonlinear meteorological and land-cover effects. Cross-validated accuracy metrics, observed-versus-predicted diagnostics, SHAP-based driver rankings, and spatial generalization to held-out stations will be reported; these outputs will assess feasibility for augmenting regulatory monitoring in data-limited metropolitan regions.

Conclusions: The proposed open-data pipeline supports scalable PM2.5 surveillance aligned with IOCEA 2026 themes on air quality and climate pollutants. Linking predictive performance with interpretable drivers informs early warning and emission-control priorities for sustainable urban development. Future work will evaluate model transferability across neighboring cities and integrate fire-season and boundary-layer variables to strengthen extreme-pollution episodes.

Keywords
air quality
PM2.5
aerosol optical depth
remote sensing
machine learning
climate pollutants
Assessment of key pollutant emissions under future scenarios of residential biomass combustion in Spain and Portugal
In Silico Toxicity Profiling of Industrial Pollutants: A Computational NAM Approach to Assessing PFAS Ecological Risks in Zebrafish (Danio rerio)