Sequential Random Split with Random Feature Subsetting for imbalanced Iron Deficiency Anemia Classification: A Machine Learning Study
Автор: S. N. Lakshmi Malluvalasa, T. Sajana
Журнал: International Journal of Engineering and Manufacturing @ijem
Статья в выпуске: 4 vol.16, 2026 года.
Бесплатный доступ
Iron Deficiency Anemia (IDA) is the most common type of anemia and can adversely affect quality of life by reducing oxygen delivery to body tissues. Accurate diagnosis is essential but remains challenging due to the manual interpretation of Complete Blood Count (CBC) parameters, which can be time-consuming and susceptible to human error. Although machine learning models have explored for early IDA detection, imbalanced datasets often lead to biased predictions and reduced performance for minority classes. Furthermore, conventional feature selection approaches may face challenges when dealing with high-dimensional medical data, potentially affecting predictive performance and model robustness. This study proposed a Sequential Random Split (SRS) framework combined with Random Feature Subsetting strategy for iron deficiency anemia (IDA) classification. The framework explores diverse feature combinations and evaluates their impact on predictive performance using a real-world imbalanced IDA dataset. The proposed approach compared with few ensemble learning methods, including Boosting, Stacking, Random Subspace, Random Patches, Extra Trees, Voting, Bagging, and Gradient Boosting. Model performance assessed using Accuracy, Precision, Recall and F1-Score metrics. The proposed SRS framework achieved the highest classification performance among the evaluated ensemble learning approaches. For class-1 (IDA), the model achieved an accuracy of 94.11%, F1-Score of 96%, and Precision of 96%. For Class-0 (non-IDA), the model achieved a Precision of 88%, F1-Score of 88%, and recall of 88%. These findings indicate that the proposed SRS framework provides strong classification performance on the imbalanced IDA dataset. The experimental results indicate that the proposed SRS framework achieved competitive performance for IDA classification on the evaluated dataset. The combination of Sequential random split and Random Feature Subsetting contributed to improved predictive performance compared with the evaluated ensemble learning approaches. These findings suggest that the proposed framework may serve as a useful machine learning strategy for supporting IDA diagnosis, while further validation on larger and more diverse clinical datasets required to assess its generalizability and robustness.
Iron Deficiency Anemia, Anemia, Classification, Machine Learning, Imbalanced Data, Ensemble learning, Diagnostics
Короткий адрес: https://sciup.org/15020575
IDR: 15020575 | DOI: 10.5815/ijem.2026.04.03