Artificial Intelligence in Natural Product Antidiabetic Drug Discovery: From Ethnopharmacological Knowledge to Machine Learning-Guided Lead Identification

Main Article Content

Jagannath Nalavade, Rajani Sajjan, Shweta Lilhare, Vishnu Suryawanshi, Amar Buchade, Nandkishor P Karlekar

Abstract

When it comes to antidiabetic agents, the core source is plants and the other natural products or their subsequent products in addition to their derived accounts for a sizable proportion of the small molecule drugs that have been allowed from 1981 to 2019. 


The traditional method of ethnobotanical examination to a categorised lead is apparently lengthy, and the chemical space of plant secondary metabolites is so huge that it has to be examined comprehensively by test. For minimising the time of the process, machine learning is the solution. The present paper studies the application of these methods throughout the antidiabetic organic product system and calculates the authenticity of the performance as an outcome. The primary test is the molecular target landscape of type two mellitus. Objectives contrast noticeably in their compliance for data-driven modelling, with carbohydrate-digesting enzymes substantially competent served by existing bioactivity data than intracellular signalling proteins.


The pipeline is then drawn through the digitally recorded data into IMPPAT and COCONUT a machine-readable atlas with the selection of molecular representation, traditional and graph-based structures reshaped by AlphaFold and predictive ADMET filtering. The recorded classifier performance throughout this literature collects in the range of 0.80 and 0.93 in the area under the receiver operating characteristic curve. Many experiments and groups have shown the anticipated hits as outputs including nervonic acid against α-glucosidase and proline-rich peptides against dipeptidyl peptidase-4. There are five methodological flaws that has been observed: tiny and narrow chemically training sets, arbitrary instead of scaffold-based data splitting, resampling applied earlier instead of within cross-validation folds, domain analysis applicably absent, and future experimental testing is limited. Overlooking these factors resulted in accuracy characterises the data partition than the underlying structure activity relationship. The proposed minimal standard is a seven-stage reporting pipeline with the argument that through the machine learning methodology in this field is more precisely described as a candidate trigger than an activity furcating, and the value should be assessed by hit rate in prospective assay than by reviewing classification accuracy. The subsequent research identifies that explainability federated learning across institutional datasets and the integration of metabolomic profiles.

Article Details

Section
Articles