Quantifying Evaluation Leakage in Bioprocess Machine Learning: A Run-Level Framework for Honest Anomaly Detection and Uncertainty-Aware Monitoring
Main Article Content
Abstract
Machine learning studies of fed-batch fermentation monitoring routinely report near-perfect anomaly detection performance, yet this rarely translates into deployed monitoring systems, because evaluation protocols leak information from the test partition into training, calibration, or threshold selection. We quantify four leakage modes on a 200-run Monod-kinetics fed-batch benchmark with four injected fault types (dissolved-oxygen crash, pH excursion, substrate overfeed, agitation failure). An identical XGBoost model evaluated under a naive pipeline (sample-level splitting, oracle Isolation Forest contamination) shows ROC-AUC inflated by 0.039, PR-AUC by 0.077, and F1 by 0.103 relative to a run-level, leakage-controlled pipeline. Under the run-level protocol, XGBoost achieves ROC-AUC = 0.9559 (95% BCa CI [0.9425, 0.9667]) on a sealed test partition of 30 runs (N = 2,910; 5.21% prevalence). Isotonic regression on a held-out calibration partition reduces Expected Calibration Error from 0.0140 to 0.0053 (ΔECE = 0.0087; 95% CI [0.0005, 0.0139]) with no detectable loss of discrimination (TOST, δ = 0.05). A feature-leakage audit shows two faults (substrate overfeed, agitation failure) achieve near-perfect AUC (≥ 0.9999) via single-feature collapse and are reported as simulation artefacts, while do_crash and pH_excursion (AUC 0.9148–0.9867) reflect genuine learned detection. Leakage also distorts interpretability: naive-pipeline SHAP rankings diverge materially from run-level, sealed-test-only rankings (Kendall's τ = 0.49; Spearman's ρ = 0.63). The same protocol, re-applied to the Tennessee Eastman Process benchmark, yields ROC-AUC = 0.8885 for a separately trained model, supporting portability of the evaluation protocol rather than of any trained model. All results derive from simulated, injection-labelled data and quantify pattern recovery within the simulation domain, not fault-detection capability in physical bioreactors. The central finding is that evaluation design changes apparent performance — and apparent feature importance — by a margin larger than differences typically attributed to model choice.
