Skip to content

Physician Reasoning on Management Cases With Large Language Models

Management Reasoning With AI Chat Bots

Status
Completed
Phases
NA
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT06208423
Enrollment
92
Registered
2024-01-17
Start date
2023-12-28
Completion date
2024-04-19
Last updated
2024-09-27

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Clinical Decision-making

Keywords

clinical decision support

Brief summary

This study will evaluate the effect of providing access to GPT-4, a large language model, compared to traditional management decision support tools on performance on case-based management reasoning tasks.

Detailed description

Artificial intelligence (AI) technologies, specifically advanced large language models like OpenAI's ChatGPT, have the potential to improve medical decision-making. Although ChatGPT-4 was not developed for its use in medical-specific applications, it has demonstrated promise in various healthcare contexts, including medical note-writing, addressing patient inquiries, and facilitating medical consultation. However, little is known about how ChatGPT augments the clinical reasoning abilities of clinicians. Clinical reasoning is a complex process involving pattern recognition, knowledge application, and probabilistic reasoning. Integrating AI tools like ChatGPT-4 into physician workflows could potentially help reduce clinician workload and decrease the likelihood of mismanagement. However, ChatGPT-4 was not developed for clinical reasoning nor has it been validated for this purpose. Further, it may be subject to disinformation, including convincing confabulations that may mislead clinicians. If clinicians misuse this tool, it may not improve reasoning and could even cause harm. Therefore, it is important to study how clinicians use large language models to augment clinical reasoning prior to routine incorporation into patient care. In this study, participants will be randomized to answer clinical management cases with or without access to ChatGPT-4. Each case has multiple components, and the participants will be asked to discuss their reasoning for each component. Answers will be graded by independent reviewers blinded to treatment assignment. A grading rubric was developed for each case by a panel of 4-7 expert discussants. Discussants independently developed a rubric for each case, and then any discrepancies were resolved through multiple rounds of discussions.

Interventions

OTHERGPT-4

OpenAI's GPT-4 large language model with chat interface.

Sponsors

Beth Israel Deaconess Medical Center
CollaboratorOTHER
University of Minnesota
CollaboratorOTHER
Stanford University
Lead SponsorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
PARALLEL
Primary purpose
TREATMENT
Masking
SINGLE (Outcomes Assessor)

Masking description

The grading of responses will be performed by assessors blinded to participant identity and treatment assignment.

Intervention model description

The trial will be designed as a randomized, two-arm, single-blind parallel group study.

Eligibility

Sex/Gender
ALL
Healthy volunteers
Yes

Inclusion criteria

* Participants must be licensed physicians and have completed at least post-graduate year 2 (PGY2) of medical training. * Training in Internal medicine, family medicine, or emergency medicine.

Exclusion criteria

* Not currently practicing clinically.

Design outcomes

Primary

MeasureTime frameDescription
Management ReasoningWithin one-hour studyPercent correct (range: 0 to 100) for each case.

Secondary

MeasureTime frameDescription
Time Spent on ManagementWithin one-hour studyTime (in minutes) participants spend per case between the two study arms.

Countries

United States

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Feb 6, 2026