Explainable Machine Learning for Predicting Antibacterial Activity of Natural Products
Main Article Content
Abstract
Natural products are a valuable source of bioactive compounds with considerable potential for antibacterial drug discovery. However, experimental screening of large natural-product collections is often time-consuming and resource-intensive. In this study, an explainable machine-learning framework was developed to predict antibacterial activity of natural products using data from the NPASS 3.0 database. Bacterial minimum inhibitory concentration (MIC) records were processed and classified using an operational threshold of ≤8 µg/mL, resulting in a final dataset of 3,130 compounds, including 624 active and 2,506 inactive compounds. Ten molecular descriptors were combined with 2,048-bit Morgan fingerprints to represent the molecular structures. Logistic Regression, Support Vector Machine, Random Forest and XGBoost models were evaluated using five-fold cross-validation and an independent test set. On the independent test set, Random Forest achieved the highest accuracy (0.8530) and ROC-AUC (0.8698), while XGBoost achieved the highest recall (0.6320), F1-score (0.6245) and MCC (0.5295). XGBoost was further interpreted using SHAP analysis to examine the contribution of molecular descriptors and fingerprint features to individual predictions. Morgan fingerprint features showed the highest overall contributions, while LogP was the most influential conventional molecular descriptor. The model was also used to prioritize natural products with high predicted probabilities of antibacterial activity. The proposed framework provides an interpretable computational approach for antibacterial activity prediction and can support the prioritization of natural products for further biological investigation.
