Skip to content

Teaching Doctors in Training to Reason With Artificial Intelligence

Teaching Internal Medicine and Family Medicine Residents to Reason With Generative AI: A Multi-Site Randomized Controlled Trial (TEACH-AI)

Status
Not yet recruiting
Phases
Unknown
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT07813754
Acronym
TEACH-AI
Enrollment
200
Registered
2026-09-10
Start date
2026-09-01
Completion date
2027-06-01
Last updated
2026-09-10

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Artificial Intelligence (AI), Clinical Decision-making, Clinical Reasoning, Medical Education

Keywords

artificial intelligence, medical education, clinical reasoning, clinical decision-making, human-computer interaction

Brief summary

The purpose of the TEACH-AI study is to assess whether a brief, structured workshop on artificial intelligence can improve the performance of medicine doctors in training (i.e. residents) in their diagnostic and management reasoning. In this multi-site randomized controlled trial, internal medicine and family medicine residents are assigned either to receive an in-person workshop on safe, effective LLM use before a standardized AI-assisted assessment, or to complete the same assessment before receiving the workshop. Residents will review clinical cases that are fully synthetic, no protected health information is used, using a password protected LLM interface.

Detailed description

Large language models (LLMs) have rapidly entered routine use in medical education and clinical practice. Prior randomized trials have shown that, although LLMs can outperform individual clinicians on some reasoning benchmarks, providing physicians with access to an LLM does not necessarily improve their diagnostic or management performance, and LLMs alone may perform better than physician-plus-LLM teams. These findings suggest that human-AI collaboration is fraught, non-trivial and may require explicit training. TEACH-AI is a pragmatic, multi-site, two-arm randomized controlled trial embedded within protected residency didactic time. The primary objective is to determine whether a single in-person workshop on the basics of LLMs and best practices of prompting/verification strategies improves residents' performance on an AI-assisted assessment, compared with residents who complete the simulation before receiving the workshop. All participants will ultimately receive the same workshop and the same assessment (either workshop-first vs assessment-first). The simulation consists of multiple fully synthetic vignettes delivered via a password protected LLM interface. Residents interact freely with the LLM using natural-language prompts and then submit structured final responses regarding aspects such as leading diagnosis, differential, management plan, and/or justification. The platform will record prompts, model outputs, final answers, and timing. Vignette scoring combines correctness of the final diagnosis or management plan. Scoring is conducted by blinded faculty using standardized rubrics and then any discrepancies will be resolved through multiple rounds of discussions. The trial will enroll up to 200 residents across four ACGME-accredited programs (internal medicine at Beth Israel Deaconess Medical Center, Stanford University, and Cambridge Health Alliance; family medicine at AdventHealth Orlando).

Interventions

OTHERWorkshop-First

Participants will attend a workshop during their didactic time that will review various aspects of large language models including fundamentals and best practices of interacting and interpreting their outputs.

Sponsors

Beth Israel Deaconess Medical Center
Lead SponsorOTHER
Stanford University
CollaboratorOTHER
Cambridge Health Alliance
CollaboratorOTHER
AdventHealth
CollaboratorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
PARALLEL
Primary purpose
DIAGNOSTIC
Masking
SINGLE (Outcomes Assessor)

Masking description

The grading of responses will be performed by assessors blinded to participant identity and whether workshop-first or assessment-first assignment.

Intervention model description

The trial will be designed as a randomized, two-arm, single-blind parallel group study.

Eligibility

Sex/Gender
ALL
Healthy volunteers
Yes

Inclusion criteria

* Participants must be licensed physicians and have started at least post-graduate year 1 (PGY1) of medical training. * Training in Internal medicine or family medicine or emergency medicine. * Able to provide informed consent and complete assessments in English.

Exclusion criteria

* Not currently practicing clinically. * Resident directly involved in the design of the study, development of the workshop, or creation/piloting of the assessment vignettes. * Resident who declines or withdraws consent.

Design outcomes

Primary

MeasureTime frameDescription
Total Score on Expert-Development RubricsWithin 24 hours of assessment completion.The primary outcome will be the number of correct responses for all cases using select questions from expert-developed scoring rubrics. The rubrics were developed using a Delphi consensus process by expert physicians as established by Goh et al (Nature Medicine, 2025; DOI: 10.1038/s41591-024-03456-y). The primary outcome will be analyzed at the case level, comparing performance between the randomized study groups, with a higher number of correct responses indicating a better outcome.

Secondary

MeasureTime frameDescription
Time Spent on ManagementWithin 24 hours of assessment completion.Comparison of time spent per case between the two study groups.
Prompt CountWithin 24 hours of assessment completion.Comparison of the number of resident-initiated prompts to the LLM per case between the two study groups.
Management Reasoning Using Expert-Derived RubricsWithin 24 hours of assessment completion.Comparison of management reasoning accuracy based on the number of correct responses per case between the two groups on expert-developed scoring rubrics. The rubrics were developed using a Delphi consensus process by expert physicians as established by Goh et al (Nature Medicine, 2025; DOI: 10.1038/s41591-024-03456-y). The primary outcome will be analyzed at the case level, comparing performance between the randomized study groups with more correct responses indicating a better outcome.
Diagnostic ReasoningWithin 24 hours of assessment completion.Comparison of diagnostic reasoning accuracy based on the number of correct responses per case between the two groups.

Countries

United States

Contacts

CONTACTJacob M Koshy, MD, MPH
jkoshy@bidmc.harvard.edu617-754-4677

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Sep 11, 2026