Effects of Generative Artificial Intelligence on Higher-Order Cognitive Outcomes in Health Science Education: A Meta-Analysis of Interventional Studies
Main Article Content
Abstract
Background: Generative artificial intelligence (GenAI) is increasingly used in health science education, but its effects on higher-order cognitive outcomes remain uncertain.
Objective: To quantify the effects of GenAI-supported educational interventions on critical thinking, problem-solving, clinical reasoning, and decision-making among health science learners.
Methods: This meta-analysis used quantitatively eligible interventional studies identified within a prospectively registered evidence-synthesis protocol (PROSPERO CRD420251128495). PubMed, Scopus, the Cochrane Library, and relevant gray literature were searched between September and November 2025. Studies were eligible for quantitative synthesis when they evaluated a GenAI-supported educational intervention and reported sufficient data to calculate standardized mean differences. Hedges’ g with 95% confidence intervals (CIs) was synthesized using DerSimonian-Laird random-effects models. Heterogeneity was assessed using Cochran’s Q, I², and τ². Publication bias and small-study effects were explored using funnel plots and Begg’s and Egger’s tests, and robustness was examined using leave-one-effect-estimate-out sensitivity analysis.
Results: Fifteen unique interventional studies contributed 21 outcome-domain effect estimates to the revised meta-analysis after exclusion of Wu et al., whose conventional AI diagnostic platform did not meet the operational definition of GenAI. The overall pooled effect favored GenAI-supported education (Hedges’ g = 0.58, 95% CI 0.22–0.94, p = 0.002), with considerable heterogeneity (Q = 475.07, p < 0.001; I² = 95.79%; τ² = 0.651). Decision-making showed a large positive pooled effect (g = 1.52, 95% CI 0.78–2.26), and critical thinking also showed a large positive effect (g = 1.75, 95% CI 1.08–2.41). Problem-solving showed a small negative pooled effect (g = −0.22, 95% CI −0.32 to −0.12), whereas clinical reasoning was not statistically significant (g = −0.19, 95% CI −0.79 to 0.41). Begg’s (p = 0.084) and Egger’s (p = 0.176) tests were not statistically significant. Leave-one-effect-estimate-out analyses produced pooled estimates ranging from 0.47 to 0.65.
Conclusions: GenAI-supported education was associated with a moderate overall improvement in higher-order cognitive outcomes, but effects differed substantially across domains. The very high heterogeneity, together with dependence among multiple effect estimates contributed by some studies, requires cautious interpretation of the overall pooled estimate. Domain-specific findings suggest potential benefits for decision-making and critical thinking, while problem-solving and clinical reasoning require more careful instructional design and further study.
