EventsThe 6th International Electronic Conference on Applied Sciences
Published
This submission belongs to the session S3. Computing and Artificial Intelligence of the event The 6th International Electronic Conference on Applied Sciences
Published date
03 Dec, 2025
Academic Editor
author-avatarLucia Billeci
Citation
Bilal Ibrahim Maijamaa, Salim Ahmad, Aminu Musa, Abdullahi Ishaq, Abida Ayuba, Ensemble-Based Imputation for Handling Missing Values in Healthcare Datasets: A Comparative Study of Machine Learning Models, in Proceedings of The 6th International Electronic Conference on Applied Sciences, 9 December–11 December 2025, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Ensemble-Based Imputation for Handling Missing Values in Healthcare Datasets: A Comparative Study of Machine Learning Models

image
1. Computer Science Department, Federal University Dutse, Dutse 720211, Nigeria, Nigeria
2. Information Technology, Federal University Dutse, Dutse 720211, Nigeria, Nigeria
Abstract

Missing values significantly impede data analysis and machine learning, especially in healthcare where complete data is vital. They can reduce predictive model performance, making robust imputation essential. Traditional methods like mean and median substitution often perform poorly with high missingness. This study compares traditional statistical imputers with machine learning models for handling missing data. Seven machine learning algorithms were tested on four datasets with substantial missing values, revealing performance declines in both statistical and ML-based imputation methods when missingness was high. To overcome this, the study proposes a stacking ensemble combining Random Forest, Linear Regression, and Ridge Regression to boost predictive accuracy and reduce error.The proposed model was evaluated using standard metrics, such as Accuracy and Root Mean Squared Error (RMSE), and was compared against individual models and traditional imputation methods. Results show that the ensemble technique achieved accuracy of 98.2% and RMSE 0.2093 outperforming all seven individual machine learning models and statistical methods on the breast cancer dataset. RF with 97.08% and XGBoost with 95.9% accuracy also consistently outperformed statistical imputers across all datasets. Notably, the Decision Tree model exhibited poor performance across all datasets, with high RMSE and low accuracy. These findings highlight the importance of selecting appropriate imputation strategies and algorithms to enhance predictive accuracy in the presence of missing data. This work contributes to the growing body of research on machine learning-based imputation and predictive modeling in healthcare and other domains.

Keywords
Breast Cancer Prediction
Machine Learning
Imputation Methods
Predictive Performance
Missing Values
Model Comparison
Data Preprocessing.
Machine Learning-Based Hybrid Model for Improved Crop Yield Correlation Analysis: A Data-Driven Assessment
Synthesis of Naphthalen-2-yl 2-thiocyanatoacetate and Its Application as a Selective Photometric Reagent for Ni(II) Detection