EventsThe 1st International Online Conference on Forecasting
Published
This submission belongs to the session S3. Forecasting and Econometric Models of the event The 1st International Online Conference on Forecasting
Published date
16 Sep, 2026
Academic Editor
author-avatarAlessia Paccagnini
Citation
Ratikorn Sornumpol, Macroeconomic Feature Extraction for S&P 500 Return Forecasting and Market Regime Classification Using PCA and Explainable Machine Learning, in Proceedings of The 1st International Online Conference on Forecasting, 21 September–22 September 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Macroeconomic Feature Extraction for S&P 500 Return Forecasting and Market Regime Classification Using PCA and Explainable Machine Learning

1. Financial Engineering, World quant university, Bangkok 10250, Thailand
Abstract

Financial market forecasting is challenging because equity returns are influenced by nonlinear macroeconomic conditions, valuation measures, and investor expectations. This study applies Principal Component Analysis (PCA) and machine learning to extract macroeconomic features for S&P 500 return forecasting and market regime classification. Monthly S&P 500 and macro-financial data are used, including price, dividends, earnings, CPI, long-term interest rate, real price, real dividends, real earnings, and CAPE/PE10. PCA is applied to reduce multicollinearity and transform correlated variables into orthogonal components. The first three principal components explain approximately 96.0% of the total variance, indicating that PCA effectively summarizes the main macroeconomic information in the dataset. Forecasting models are evaluated using RMSE and , while regime classification models are evaluated using accuracy and score. The empirical results show that PCA slightly improves selected models, particularly Linear Regression, Random Forest, and XGBoost. However, all return forecasting models produce negative out-of-sample values, suggesting that monthly S&P 500 returns remain difficult to predict using macroeconomic variables alone. For regime classification, PCA improves Logistic Regression slightly, but overall classification performance remains modest. The findings suggest that PCA is useful for reducing feature redundancy and improving model stability, although additional volatility, momentum, and rolling macroeconomic features are needed for stronger predictive performance.

Keywords
Principal component analysis
machine learning
S&P 500
macroeconomic indicators
return forecasting
market regime classification
Poster
IOCFC_2026_Macroeconomic_Feature_Extraction_Poster_EN.pdf
Specialized versus Generic XGBoost Models for Direct Multi-Step Time Series Forecasting: A Comparative Study on the ETT Benchmark
High-Frequency Tourism Flow Forecasting in a Historical City Center: A Rolling Origin Evaluation of DHR, Splines, and TBATS Models