Machine Learning and Hybrid Ensemble Models for Maternal Health Risk Prediction
Volume 3, Issue 2, January 2026, Pages 22-32
https://doi.org/10.48308/jicse.2026.242793.1097
MohammadReza Einollahi Asgarabad, Mahsa Akhbari
Abstract This study presents an applied, multi-stage machine-learning pipeline for maternal health risk classification using the Maternal Health Risk dataset (1,014 samples, six clinical features, and three risk classes). Six classifiers were evaluated across four preprocessing stages: training-fold standardization alone, Edited Nearest Neighbors (ENN)-based cleaning, Synthetic Minority Oversampling Technique (SMOTE)-based class balancing, and sequential ENN-SMOTE preprocessing. Within each fold of 10-fold cross-validation, the scaler, ENN, SMOTE, model fitting, and ensemble construction were performed using only the training fold; the corresponding test fold remained untouched until evaluation. A fixed majority-voting ensemble comprising Random Forest, Extra Trees, and XGBoost was evaluated using class-wise, macro-averaged, and weighted-averaged accuracy, precision, and F1-score. In Stage 4, the hybrid ensemble achieved a weighted accuracy of 91.5% ± 3%, precision of 88% ± 3%, and F1-score of 87.2% ± 3%. The ensemble's precision for the high-risk class exceeded 90% across all four stages. These findings indicate competitive predictive performance on this dataset; however, direct comparison with previous studies is limited by differences in preprocessing methods, validation protocols, and model configurations. External and prospective validation is required before clinical use.

