Breast Cancer Classification from Cytomorphometric Features: A Leakage-Controlled Explainable Machine Learning Framework with Decision Curve Analysis
Main Article Content
Abstract
Background: Certain metrics can be used to artificially inflate reported performance for medical machine learning models. This can happen, for example, with preprocessing leakage. Discrimination alone does not guarantee decision utility or probability reliability.
Objective: This study assesses a leakage-controlled, calibrated, and globally explainable machine-learning framework for breast cancer cytology classification.
Methods: The Wisconsin Diagnostic Breast Cancer dataset (569 observations; 212 malignant and 357 benign; 30 numerical predictors) was evaluated using Gradient Boosting, HistGradientBoosting, and Random Forest. Repeated stratified five-fold cross-validation with 10 repeats generated 50 outer test folds. Imputation, threshold selection, and sigmoid or isotonic recalibration were confined to training data. Performance was summarized from pooled patient-level out-of-fold predictions with 2,000 stratified bootstrap resamples. Held-out permutation importance quantified global predictor relevance, and decision curve analysis assessed potential net benefit.
Results: HistGradientBoosting achieved the highest pooled PR-AUC (0.9915; 95% CI, 0.9847–0.9969) and ROC-AUC (0.9933; 95% CI, 0.9874–0.9978), with sensitivity of 0.9670, specificity of 0.9664, Brier score of 0.0228, and ECE of 0.0135. Its PR-AUC exceeded Gradient Boosting (paired-bootstrap p = 0.013), whereas the ROC-AUC difference was not statistically supported (p = 0.060). HistGradientBoosting also outperformed Random Forest for ROC-AUC and PR-AUC (both p = 0.020). Worst area, worst concave points, worst texture, area error, and worst perimeter ranked highest by held-out permutation importance. Sigmoid-based decision curves showed higher net benefit than both default strategies across thresholds of 0.01–0.60 for Gradient Boosting and HistGradientBoosting.
Conclusion: Leakage-controlled repeated cross-validation produced strong internal benchmark performance and potential decision utility. These findings require external validation and prospective clinical-impact assessment before clinical use.
