Smart Web-Based Representation of Human Health Data: Publicly Available Datasets, Contemporary Methodologies and a Comparative Analysis of Models

Main Article Content

Rohit Yadav; Reena Hooda; Manish Gupta; Mamta Yadav; Parshant Bansal

Abstract

The web is now the surface on which most human health information is aggregated, interpreted and acted upon. Hospital portals, national biobank workbenches, consumer wearable dashboards and clinician-facing decision support tools are all, architecturally, web systems that must turn heterogeneous biomedical evidence into something a person can reason about in seconds. This paper surveys that problem as it stands in mid-2026. We propose a six-layer reference architecture that separates acquisition, interoperability, representation learning, inference, presentation and governance, and we argue that the representation and presentation layers are the ones the literature has left under-specified. We then catalogue the publicly available human-health datasets that can legitimately support such systems, reporting verified scale figures and access conditions for critical-care records (MIMIC-IV v3.1: 364,627 individuals, 546,028 hospitalisations, 94,458 ICU stays), radiographic corpora (MIMIC-CXR: 377,110 images; CheXpert: 224,316 radiographs), population genomic resources (All of Us: more than 535,000 whole genome sequences; UK Biobank: whole-genome sequencing of 490,640 participants) and few-shot benchmark suites such as EHRSHOT. Nine model families are reviewed, from gradient-boosted trees through EHR transformers, clinical and general-purpose large language models, imaging foundation models, multimodal fusion, graph neural networks, retrieval-augmented generation and federated learning. A quantitative synthesis then pairs published benchmark  with the properties that decide whether a model can be served inside a web representation layer: latency class, memory footprint, interpretability, data-governance posture and regulatory exposure. Three findings stand out. Leaderboard accuracy has decoupled from deployability: reported MedQA accuracy climbed from 67.6% to 96.0% between late 2022 and late 2024, yet a 2026 Nature Medicine evaluation found specialised clinical AI tools performing no better than a general web search summary on real physician queries, with frontier general-purpose models ahead of both. Second, four of the eight quantified public resources with a defined coverage window end nine or more years before 2026, so learned representations encode historical rather than current practice. Third, of the 60 works cited here, only five address the presentation layer, against 15 each for acquisition, inference and governance. The paper closes with a four-tier evaluation protocol and eleven open challenges, calibrated to the compliance timeline the EU Artificial Intelligence Act imposes through 2027.

Article Details

Section
Articles