Ensemble Machine Learning for Student Performance Prediction in Higher Education: A Critical Review of Predictive Gains, Methodological Quality and Deployment Constraints

Okwedi Kelicha *

Computer Science Department, Faculty of Computing, Madonna University, Nigeria.

Ugochukwu Febechi Blessing

Public Health Department, Faculty of Allied Health Sciences, Enugu State University of Science and Technology, Enugu, Nigeria.

Adanna Anyanwu

Computer Science Department, Maranatha University, Lagos, Nigeria.

Chukwueke Nwagbara

Computer Science Department, Federal University of Technology Owerri, Imo State, Nigeria.

*Author to whom correspondence should be addressed.


Abstract

Ensemble machine learning has become the default modelling strategy in research on the prediction of student academic performance in higher education, and published comparisons routinely report that bagged, boosted and stacked combinations of base classifiers outperform single learners. The accumulated literature nevertheless remains difficult to interpret, because reported gains derive from heterogeneous institutional datasets, inconsistent evaluation protocols and outcome definitions that range from continuous grade estimation to binary at-risk flagging. This review critically appraises the evidence on ensemble approaches to student performance prediction in tertiary settings, with attention to the conditions under which ensembling delivers genuine improvement rather than apparent improvement generated by evaluation design. Peer-reviewed literature was identified through structured searching of five scholarly databases and indexes, supplemented by backward and forward citation searching, with selection based on relevance, methodological adequacy and contribution to the review question rather than citation volume alone. Five themes structure the synthesis: the mechanisms by which ensembles reduce predictive error in educational data; the construction of predictor sets and the trade-off between earliness and accuracy; the methodological adequacy of the evidence base, including resampling practice, metric selection and leakage risk; the transferability of models across courses, cohorts and institutions; and the translation of predictions into interpretable, equitable and actionable institutional practice. The evidence supports a modest and context-dependent ensemble advantage over well-tuned single learners, most reliably for heterogeneous tabular predictors and moderate sample sizes, and least reliably where evaluation protocols are weak. Reported accuracies in the upper ninetieth percentile are frequently associated with design features that inflate apparent performance. Confidence in the field's central claims is limited by single-institution samples, inconsistent reporting and a scarcity of evidence that predictions change student outcomes. Priorities include multi-institution external validation, prospective evaluation of intervention effects, standardised reporting of temporal validation, and fairness auditing across the full deployment cycle.

Keywords: Ensemble learning, educational data mining, learning analytics, academic performance prediction, model generalisability, algorithmic fairness, higher education


How to Cite

Kelicha, Okwedi, Ugochukwu Febechi Blessing, Adanna Anyanwu, and Chukwueke Nwagbara. 2026. “Ensemble Machine Learning for Student Performance Prediction in Higher Education: A Critical Review of Predictive Gains, Methodological Quality and Deployment Constraints”. Asian Journal of Research in Computer Science 19 (9):133-53. https://doi.org/10.9734/ajrcos/2026/v19i9911.

Downloads

Download data is not yet available.