Publicación

Integration of Classification, Survival, and Explainability Models for Breast Cancer Recurrence Prediction Using Clinical and Molecular Data

Chumbimuni, Celia · Ravines, Mariana · Huanca, Claudia · Camarena, Leonardo · Tovar, Nhikori

Resumen

Breast cancer is the most common neoplasm in the Americas, and its incidence is projected to increase by 40 % by 2040, representing a growing global clinical burden. Despite advances in diagnosis and treatment, approximately 30 % of patients experience recurrence after initial therapy. To address this challenge, we developed an integrated framework for recurrence prediction that combines classification, survival, and explainability models, using the Combined MSK Breast Cancer Cohort, derived from three public studies of the Memorial Sloan Kettering Cancer Center. The dataset included clinical and molecular variables preprocessed through KNN imputation, standardization, class balancing with SMOTE, and strict data leakage control. Six classifiers (Random Forest, XGBoost, LightGBM, Gradient Boosting, SVM, and Logistic Regression) and two survival models (CoxPH and Random Survival Forest) were trained using stratified 5-fold cross-validation. Gradient Boosting achieved the best performance (AUC=0.977; F 1=0.96), while Random Survival Forest obtained the highest concordance index (C -index =0.835), outperforming the Cox model. SHAP and LIME values identified endocrine therapy, prior recurrence, and tumor laterality as key predictors, reflecting strong clinical consistency. The proposed approach demonstrated high accuracy, transparency, and reproducibility, establishing itself as a potential tool for risk stratification and therapeutic decision support in breast cancer patients.

Autores y colaboradores

Authors

Chumbimuni, Celia
Ravines, Mariana
Huanca, Claudia
Camarena, Leonardo
Tovar, Nhikori