EventsThe 8th International Electronic Conference on Atmospheric Sciences
Published
This submission belongs to the session S7. Atmospheric Techniques, Instruments and Modeling of the event The 8th International Electronic Conference on Atmospheric Sciences
Published date
09 Oct, 2026
Academic Editor
author-avatarChun-Ho Liu
Citation
Georgios Spyrou, Serafim Kontos, Ioannis Stergiou, Paraskevi Tzoumaka, Apostolos Kelessis, Dimitrios Melas, Bias Correction of the CAMx Air Quality Model Using Deep Learning: A Comparative Evaluation of Station-Specific and Multi-Station Training Approaches, in Proceedings of The 8th International Electronic Conference on Atmospheric Sciences, 14 October–16 October 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Bias Correction of the CAMx Air Quality Model Using Deep Learning: A Comparative Evaluation of Station-Specific and Multi-Station Training Approaches

image
image
Paraskevi Tzoumaka 4
image
1. Laboratory of Atmospheric Physics, School of Physics, Aristotle University of Thessaloniki, 54124 Thessaloniki, Greece
2. Center for Interdisciplinary Research and Innovation, Aristotle University of Thessaloniki, 57001 Thessaloniki, Greece
3. Department of Mechanical Engineering, University of Western Macedonia, 50132 Kozani, Greece
4. Department of Environment, Municipality of Thessaloniki, Kleanthous 18, 54642, Thessaloniki, Greece
Abstract

Introduction

Numerical air quality models such as CAMx are essential tools for operational forecasting, yet they often exhibit systematic biases relative to ground-level observations. Machine learning techniques offer a promising approach for post-processing bias correction. However, the choice of training strategy, whether to train models independently at each monitoring site or to pool data across multiple stations, remains an open question, particularly under data-limited conditions.

Methods

In this study, two bias correction approaches were developed and compared for the CAMx model over the city of Thessaloniki, Greece, using hourly data from 2019. The first approach trains separate models for each of the eight monitoring stations (station-specific), while the second merges data from all stations into a single unified model (multi-station). Three deep learning architectures were evaluated: a hybrid LSTM-CNN, a Bidirectional LSTM (BiLSTM), and a Transformer. Input features included CAMx pollutant concentrations, WRF meteorological variables, cyclical temporal encodings, and lag features. Predictions were generated for four pollutants: NO2, O3, PM10 and PM2.5. Model performance was assessed using RMSE and the Pearson correlation coefficient (R).

Results

Both approaches substantially reduced CAMx biases, with RMSE reductions of 62–77% and Pearson R improvements of 96–175% relative to raw CAMx output. The multi-station approach outperformed the station-specific approach in 87–100% of cases, depending on the architecture. The Transformer benefited most from the multi-station strategy, achieving a 31–44% additional RMSE reduction compared to its station-specific counterpart, which is expected given that Transformer architectures typically require larger volumes of training data.

Conclusions

Pooling multi-station data for bias correction effectively overcomes the limitation of restricted training data volume and enables deep learning models to capture broader spatiotemporal patterns, offering a robust strategy for improving operational air quality forecasts.

Keywords
bias correction
air quality modeling
CAMx
deep learning
LSTM
BiLSTM
Transformer
statistical metrics
data pooling
forecast
Instrumentation and Modeling for Adaptive Wavefront Restoration in Atmospheric Turbulence Using Deformable Mirror Feedback Control
Resolving the photochemical complexity of trace gases in transitional environments using multi-directional MAX-DOAS