Machine Learning and Hybrid Ensemble Models for Maternal Health Risk Prediction

Document Type : Original Article

Authors

Department of Biomedical Engineering, SR.C, Islamic Azad University, Tehran, Iran

Abstract
This study presents an applied, multi-stage machine-learning pipeline for maternal health risk classification using the Maternal Health Risk dataset (1,014 samples, six clinical features, and three risk classes). Six classifiers were evaluated across four preprocessing stages: training-fold standardization alone, Edited Nearest Neighbors (ENN)-based cleaning, Synthetic Minority Oversampling Technique (SMOTE)-based class balancing, and sequential ENN-SMOTE preprocessing. Within each fold of 10-fold cross-validation, the scaler, ENN, SMOTE, model fitting, and ensemble construction were performed using only the training fold; the corresponding test fold remained untouched until evaluation. A fixed majority-voting ensemble comprising Random Forest, Extra Trees, and XGBoost was evaluated using class-wise, macro-averaged, and weighted-averaged accuracy, precision, and F1-score. In Stage 4, the hybrid ensemble achieved a weighted accuracy of 91.5% ± 3%, precision of 88% ± 3%, and F1-score of 87.2% ± 3%. The ensemble's precision for the high-risk class exceeded 90% across all four stages. These findings indicate competitive predictive performance on this dataset; however, direct comparison with previous studies is limited by differences in preprocessing methods, validation protocols, and model configurations. External and prospective validation is required before clinical use.

Keywords

Subjects

Volume 3, Issue 2 - Serial Number 2
January 2026
Pages 22-32

  • Receive Date 05 December 2025
  • Revise Date 05 August 2026
  • Accept Date 15 August 2026
  • First Publish Date 15 August 2026
  • Publish Date 01 January 2026