Producción científica
Permanent URI for this communityhttps://cris.pucp.edu.pe/handle/123456789/20173
Browse
2 results
Search Results
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Factors associated with working in remote areas among physicians: a machine learning approach Factores asociados al trabajo en áreas remotas en el profesional médico: un enfoque de modelos de aprendizaje automático(Universidad Nacional Mayor de San Marcos, Facultad de Medicina, 2026-01-01)Introduction. The global shortage of physicians represents a major public health challenge, with significant disparities in Peru. Objective. To determine the machine learning model that identifies the factors associated with physicians working in remote areas, defined as remote or border zones (ZAF) by the Peruvian government. Method. A quantitative, non-experimental, cross-sectional, and correlational study was conducted. A population of 10,228 primary care physicians was analyzed using the INFORHUS database of the Peruvian Ministry of Health as of July 2022. Four machine learning techniques were applied, and their performance was evaluated using the F1-Score, Matthews correlation coefficient (MCC), and geometric mean (GM) metrics to address data imbalances. Results. The Random Forest model proved to be the most effective. The most relevant variables, with the highest percentage of association with medical professionals working in remote areas, were remuneration (40.6%), age (25.6%), whether the contract type was SERUMS (14.2%), and sex (5.2%). Conclusions. The Random Forest model identified the factors associated with physicians working in remote areas (ZAF) in Peru, highlighting the predominance of economic and personal factors. This finding provides evidence to inform policies aimed at improving the distribution of physicians in Peru. - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Performance of evaluation metrics for classification in imbalanced data(Springer Science and Business Media Deutschland GmbH, 2025-03-01)This paper investigates the effectiveness of various metrics for selecting the adequate model for binary classification when data is imbalanced. Through an extensive simulation study involving 12 commonly used metrics of classification, our findings indicate that the Matthews Correlation Coefficient, G-Mean, and Cohen’s kappa consistently yield favorable performance. Conversely, the area under the curve and Accuracy metrics demonstrate poor performance across all studied scenarios, while other seven metrics exhibit varying degrees of effectiveness in specific scenarios. Furthermore, we discuss a practical application in the financial area, which confirms the robust performance of these metrics in facilitating model selection among alternative link functions.2
