Agentic Foundation Models for Explainable Multimodal Clinical Intelligence: A Human-Centered Framework for Early Disease Diagnosis and Personalized Treatment Planning

Main Article Content

Ganesh Dagadu Puri, Sachin Arun Thanekar, Renuka Sandeep Gound, Ashwini Shinde, Mahesh Ashok Bhandari, Khushal Khairnar

Abstract

Diagnostic work rarely rests on one kind of evidence. A clinician assembles the history, the imaging, the laboratory panel and, increasingly, genomic results, and does so under time pressure. Large language models perform well on each of these in isolation, yet most deployed systems still reason over a single modality, produce explanations after the fact rather than from the evidence actually used, and give the clinician no route to contest an output. This paper describes a framework that treats diagnosis as a coordination problem rather than a modeling one. Perception agents handle text, imaging and structured data separately; an orchestrator decomposes the task and routes sub-tasks among specialist agents; and a distinct explanation agent is queried after each recommendation to retrieve the inputs that drove it. Because that agent runs independently of the diagnostic pathway, its output is less likely to be a rationalization of an answer already reached. We evaluate on a pilot cohort of three cases drawn from the dataset comparing against a zero-shot baseline and a text-only specialist model. The framework improved on both across accuracy, F1-score, explanation faithfulness and clinician-rated trust, at an added latency of roughly five seconds per case. These are feasibility results from a small cohort, and we set out the prospective validation that would be needed before any claim of clinical readiness.

Article Details

Section
Articles