Introduction. Predicting the response ratio of Global Warming Potential (GWP) under biochar application is intrinsically difficult because soil processes, climate conditions, and management practices vary considerably among experimental sites. A key challenge is the trade-off between model complexity and data availability, which can reduce sample size and compromise predictive reliability when too many variables are included.
Methods. This study proposes a machine learning framework to estimate GWP response using a global dataset compiled from hundreds of studies across varied agroecosystems. The modeling pipeline is built on a reduced, data-consistent feature space of five predictors: biochar application rate, initial soil organic carbon (SOC), initial soil pH, climate zone, and treatment type. This subset was selected based on its high data completeness and consistent relevance across response variables, allowing robust model training without substantial loss of observations. Linear, regularized, and tree-based models were trained with systematic hyperparameter optimization. Adjusted R² was adopted as the performance criterion to account for differences in model complexity.
Results. Models built on this constrained feature set produced steady predictive performance across the dataset. Expanding the feature space to include additional variables, such as pyrolysis temperature or detailed environmental descriptors, reduced the available sample size while providing only modest accuracy improvements.
Conclusions. A small, physically meaningful set of predictors can reliably forecast GWP response to biochar application. The results confirm that GWP predictions are sensitive to data sparsity and that careful feature selection is critical when working with heterogeneous global datasets. The proposed framework supports scalable applications in sustainable soil management and climate impact assessment.