Diagnosis, Vignette of Fictional Patients
Conditions
Keywords
LLM, Quality of Care, Clinical Reasoning, Primary Care
Brief summary
A parallel group randomized controlled trial using a superiority framework. Clinical vignettes will be used to assess the impact of a large language model on the clinical reasoning of physicians. Quantitative analyses will be performed on graded vignette responses.
Detailed description
This study is a multi-country, parallel-group randomized controlled trial designed to evaluate whether access to a large language model (LLM) improves physician clinical decision-making. The trial uses a superiority framework and compares physicians randomized to either complete standardized clinical vignettes with access to GPT-4o or without any AI assistance. Clinical vignettes simulate common primary care conditions such as cardiovascular, respiratory, musculoskeletal, fatigue-related, and infectious diseases. Each vignette includes multiple steps in the clinical reasoning process, from initial history-taking to diagnosis, treatment, and follow-up. Physician responses are graded using rubrics developed from evidence-based, context-specific best-practice guidelines. The study is conducted across three countries-Indonesia, Kenya, and the Netherlands-representing different income levels and health system contexts. The primary outcome is performance on clinical vignettes, defined as adherence to best-practice guidelines. Secondary objectives include examining cross-country variation in physician performance, variation in performance distributions, and the role of engagement with the LLM in shaping outcomes.
Interventions
GPT-4o provided via an iFrame in the online Qualtrics environment
Sponsors
Study design
Intervention model description
We designed clinical vignettes to simulate patient-physician consultations. We follow the structure of clinical performance and value vignettes, a vignette design validated to perform well against standard patient designs. Our vignette design follows 9 stages: (1) Presenting Problem & Initial Differential Diagnosis, (2) Asking about patient history, (3) Additional Differential Diagnosis, (4) Listing Physical Exams, (5) Differential Diagnosis, (6) Additional diagnostic (Lab) Tests, (7) Final Differential Diagnosis, (8) Medication, (9) Follow-Up/Advice. Participants were randomly assigned either to the control or intervention group using simple randomization and asked to complete the vignettes in an online environment. Intervention group participants were given access to an LLM (GPT-4o)
Eligibility
Inclusion criteria
* Registered medical physicians * Training in internal or family medicine
Exclusion criteria
* Not currently practicing clinically
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Percentage Correct Score | During Evaluation | Following Peabody et al (2000), the primary outcome is a percentage correct score across all steps in a vignette. This is generated by dividing the weighted total sum of rubric items assessed as present by the total number of rubric items possible in a vignette. Rubric items will be weighted with regards to their relevance by our expert panel. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Quality Per Answer | During Evaluation | This outcome is generated as the average weight of rubric items assessed as present across vignettes. As each item is provided a weight (0.33,0.5,1), the average weight is the sum of weights divided by the number of answers marked as present. |
| Number of Answers | During evaluation | This outcome is generated as the count of the total number of answers assessed as present by reviewers per vignette |
| Less obvious answers | During evaluation | This outcome is generated as the number of answers given that are less obvious, i.e. mentioned less frequently by the control group. If the answer is mentioned by 25% or less of the control group, it is considered less obvious. |
Countries
Indonesia, Kenya, Netherlands
Contacts
Maastricht University