Predicting Emergency Department Mortality Risk Using Machine Learning Algorithms The Case of Yekatit 12 Hospital Medical College
No Thumbnail Available
Date
2025
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Addis Ababa University
Abstract
The increasing burden of emergency department (ED) mortality in resource-constrained settings, such as Ethiopia, underscores the need for innovative approaches to enhance clinical decision-making. Though its use in low- and middle-income countries (LMICs) is still relatively unexplored, machine learning (ML) holds great promise for improving mortality risk prediction using electronic medical record (EMR) data. While machine learning (ML) has been widely used in high-income healthcare settings, little is known about how well it works for predicting mortality risk in LMIC emergency departments (EDs), especially when using locally relevant datasets and overcoming issues like data imbalance and unstructured EMR data. The goal of the study is to identify important predictive variables, evaluate the effectiveness of different machine learning models in a resource-constrained setting, and develop and assess ML-based model for predicting mortality risk among emergency department patients
at Yekatit 12 Hospital Medical College in Addis Ababa, Ethiopia. A two-year dataset of EMRs from Yekatit 12 Hospital (October 2022–September 2024, or
2015–2016 in Ethiopian calendar) was used for a retrospective analysis. Attribute selection, Data cleaning, Text mining, Data transformation, and Feature engineering were all used in the methodology to glean insights from unstructured data. Four machine learning models: Random Forest, k-Nearest Neighbors, Support Vector Machines, and XG Boosting - were trained and evaluated. Undersampling, hybrid approaches, and the Synthetic Minority Oversampling Technique (SMOTE) were used to rectify the data imbalance. Evaluation parameters such as accuracy, precision, recall, F1-score, and Area Under the Receiver Operating Characteristic Curve (AUC-ROC) were used to evaluate the model's performance. To improve model robustness, RandomizedSearchCV was used for hyperparameter optimization. The Random Forest model, optimized with RandomizedSearchCV, performed better than other models, with an F1-score of 0.68, precision of 0.90, recall of 0.55, and AUC-ROC of 0.81. Patients‘ clinical diagnoses, age, vital signs, and some laboratory result were critical predictive variables. The use of SMOTE and hybrid sampling techniques greatly enhanced
IV model performance on imbalanced data, with Random Forest showing superior generalization across evaluation metrics. This study shows that machine learning (ML), specifically Random Forest with hyperparameter tuning, can provide a scalable tool for clinical decision support by accurately predicting the risk of ED mortality in an LMIC setting using routine EMR data. In order to improve predictive accuracy, the results emphasize how critical it is to address data imbalance and use text mining for unstructured data.
Description
Keywords
Electronic Medical Record Emr Predicting Mortality Risk Machine Learning