Publicación

Data Mining to Identify University Student Dropout Factors

Yuri Reina Marín · Lenin Quiñones Huatangari · Omer Cruz · Jorge Luis Maicelo Guevara · Judith Nathaly Alva Tuesta · Einstein Sánchez Bardales · River Chávez Santos

Resumen

University dropout poses academic, social, and economic challenges that call for effective prevention strategies. The objective was to identify determining factors of student dropout through educational data mining and machine learning models. A survey was administered to 527 undergraduate students, and the data were processed with classification algorithms (Adaboost, Gradient Boosting, Extra Trees, Random Forest, Decision Tree, and XGBoost), complemented with interpretation techniques such as SHAP and sensitivity analysis. The results revealed that, in addition to prior academic performance (GPA), psychological support emerged as the most influential predictor across all models, followed by institutional and socioeconomic variables, including academic program, age, and parental job stability. Integrating psychological, institutional, and family factors into predictive systems enhances model accuracy and provides practical evidence to inform educational policies, strengthen student support programs, and design early interventions to promote retention in higher education.

Autores y colaboradores

Authors

Yuri Reina Marín
Lenin Quiñones Huatangari
Omer Cruz
Jorge Luis Maicelo Guevara
Judith Nathaly Alva Tuesta
Einstein Sánchez Bardales
River Chávez Santos

Palabras clave

Academic performance Educational data mining Higher education Machine learning Student retention University dropout