TrustXAI-Derm: An Empirical Audit of Explanation Reliability, Failure Prediction, and Deferral in Dermoscopic AI
Main Article Content
Abstract
Background: When applied to the skin, the typical result provided by dermoscopic AI models is aggregate accuracy, without considerations of accuracy reliability or the reliability of visual explanations across diagnostic classes, model architectures, lesion sizes, or estimated skin tones.
Objective: We present an empirical audit (not a new classifier) of the reliability of dermoscopic explanation on these dimensions.
Methods: Given solely the mask overlap as explanation ground truth, we define a mask-overlap Trustworthy Index for Explainable AI (TIxAI) and test it per diagnostic class (GAP-1), train a meta-model to predict an explanation failure based on embeddings and Monte Carlo Dropout (MC-Dropout) uncertainty prior to computing any explanation (GAP-2), audit explanation quality and uncertainty across an Individual Typology Angle (ITA) skin-tone proxy (GAP-3), and prototype a Composite Risk Score for clinician deferral (GAP-4).
Results: Reliability of explanations, when tested on a split of the patients (n = 990) was significantly different across the diagnostic classes (Kruskal–Wallis H = 30.65, p = 2.95×10⁻⁵) as well as with lesion size (H = 269.5, p = 3.0×10⁻⁵⁹) with the latter class showing the largest effect. Compared to the two, ConvNeXt-Tiny achieved better balanced accuracy (0.763 vs. 0.689), but significantly lower calibration (0.308 vs. 0.074) (Expected Calibration Error) ECE. The best result was the failure-prediction meta-model (5-fold cross-validated AUC-ROC 0.832, 95% BCa CI 0.793–0.866). The difference in MC-Dropout entropy was significant across ITA buckets (p = 0.0004), but not for TIxAI (p = 0.20), so we do not consider this an evidence of fairness in the latter case, but rather an underpowered null. The following candidate trust signals were null results: Grad-CAM++/SHAP disagreement, mask-free saliency geometry, and contrastive analysis. Unresolved issues are highlighted instead of being hidden, such as reported pre-convergence (epoch 60/150), or reported near-zero median TIxAI (0.0006) as a possible artifact of a Grad-CAM++ target-layer configuration, but not as a genuine result, and the single-seed result of DenseNet121's raw-versus-balanced accuracy inversion is not diagnosed. A prototype deferral score focused residual errors in the deferred subset, and is based on a single split and needs to be validated using a bootstrap method. External validation and a dermatologist reader study are not yet performed.
Conclusion: We thus offer the rigorous empirical study as a self-audited work where the contribution is the audit approach and its positive, null and unresolved results, and not a clinical tool that can be deployed. Clinical/Industrial Impact: TrustXAI-Derm proposes an auditing path towards explanations that are presented to clinicians for decisions, highlighting those explanations that are low-trust before the system has arrived, and proposes a prototype of such an explanation that concentrates the residual error in the explanations withheld upon deployment.
