A Bibliometric and Systematic Review of Machine Learning Applications for Predicting Student Performance in Higher Education: Insights via the TCCM Framework
Main Article Content
Abstract
Background: Higher education institutions collect large amounts of student data through learning management systems, student information systems, online assessments and academic records. These data sources have increased the use of machine learning for predicting student performance. Although this area has grown in recent years, the literature remains scattered across education, computer science and learning analytics.
Purpose: This study presents a bibliometric and systematic review of machine learning applications for predicting student performance in higher education. The review also uses the TCCM framework to examine the literature through four dimensions: Theory, Context, Characteristics and Methodology.
Methods: The study followed the PRISMA framework. Data were collected from the Scopus database. The search covered peer-reviewed English-language publications from 2017 to 2025. The search produced 1,193 records. After duplicate removal and screening, 1,156 records were included in the final review. Bibliometric analysis was conducted using R software, Bibliometrix and Biblioshiny.
Results: The results of this study have shown an upward trend in machine learning based on the number of articles related to predicting students' performance. The 47.69% Annual Growth Rate (AGR) for machine learning shows there has been an increased interest in using Machine Learning methods to predict how well a student will do in school. Most of these studies are found in specific areas around the world; and at the institution level. Student data that can be used as predictors includes: attendance, assignments completed by student; log files from Learning Management Systems (LMS); test scores, quizzes etc.; prior academic history/achievement. Many researchers are now utilizing classification algorithms and ensemble models. A large number of studies rely upon data collected from one single institution. Limited studies utilized external validations, theoretical bases to select variables, explanations for their model's predictions; and actual use within the classroom to evaluate if their model could make valid predictions about students' potential success.
Conclusion: The study indicates that machine learning is a major methodology in using data from students to predict student success at the post-secondary level. While there are many studies in this area; most of these focus on technical aspects of model building rather than providing insight into the meaning and relevance of their findings. Therefore, future studies will need to develop models that are based on well-established theories of education as well as clearly define constructs used to build those models. Additionally, researchers will need to validate the generalizability of models by testing them across different contexts and provide clear descriptions of how they validated their results.
