Skip to content

Comparing Artificial Intelligence and Physicians: A Vignette-Based Study in Pediatric Clinical Decision-Making

A Prospective, Cross-Sectional, Vignette-Based Observational Study Comparing Clinical Decision-Making Performance of Pediatriciansand AI Models

Status
Completed
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07179861
Enrollment
30
Registered
2025-09-18
Start date
2025-08-27
Completion date
2025-09-11
Last updated
2025-09-23

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Artificial Intelligence (AI) in Diagnosis, Clinical Decision-making, Decision Support Systems, Clinical, Pediatrics

Brief summary

This study evaluates how well anonymized artificial-intelligence (AI) tools perform on standardized pediatric case vignettes and whether showing AI suggestions can improve clinicians' answers. About 30 board-certified/eligible pediatric specialists at a single hospital complete a one-time session. Participants are randomized to two groups. Group A (n≈15): physicians answer each vignette once. Group B (n≈15): physicians answer and rate confidence (1-10), then review anonymized suggestions from five different AI tools (tool names not shown) and may keep or change their answer; changes and confidence are recorded. Primary focus: measure AI performance (diagnostic accuracy, medication-dosing accuracy, interpretation accuracy) overall and by difficulty tier, and record AI response time. Secondary focus: quantify how AI suggestions affect human performance (change in accuracy, direction of change, confidence shift, and time). No patients or biospecimens are involved; risks are minimal (time and possible discomfort with performance review). Findings may inform safe, evidence-based ways to use AI alongside clinicians in pediatrics.

Interventions

OTHERAI Suggestions (Anonymized 5-tool panel)

What: Display of AI-generated suggestions for each vignette, aggregated from five large language model tools (names not shown to participants). When/Who: Shown only in Group 2, after the physician's initial answer and confidence score. Purpose: Measure AI performance (primary) and quantify the effect of AI suggestions on physicians' answers (secondary). Applies to: Group 2.

OTHERConfidence Rating Task (1-10 Likert)

What: Self-rated confidence for the initial answer on a 1-10 scale. When/Who: Group 2 before viewing AI suggestions. Purpose: Quantify confidence changes pre- vs post-AI and relate confidence to correctness. Applies to: Group 2.

Sponsors

Haseki Training and Research Hospital
Lead SponsorOTHER

Study design

Observational model
COHORT
Time perspective
CROSS_SECTIONAL

Eligibility

Sex/Gender
ALL
Age
28 Years to 40 Years
Healthy volunteers
Yes

Inclusion criteria

* Board-certified or board-eligible pediatric specialist (general pediatrics) (in the first 10 years of expertise) * Actively practicing at the participating institution/network at the time of enrollment. * Able and willing to complete all vignette items individually in a single session and to follow study instructions for the assigned cohort (direct answers or confidence rating + viewing anonymized AI suggestions). * Fluent in Turkish and able to use a computer interface. * Provides written informed consent.

Exclusion criteria

* Pediatric subspecialist practice as primary role (e.g., cardiology, infectious diseases, neurology, neonatology, etc.), to maintain a homogeneous general pediatrics cohort. * Prior access to or participation in creating the study vignettes, answer keys, or scoring rubrics; direct involvement with the study team. * Inability to complete the session without external help or use of non-protocol resources (internet/AI tools) during answering (outside of anonymized AI suggestions shown by the system in Group 2). * Failure to complete ≥90% of items or major protocol deviation (e.g., discussion with others during the task). * Any condition judged by investigators to interfere with valid participation (e.g., severe time constraints, inability to provide consent).

Design outcomes

Primary

MeasureTime frameDescription
AI Interpretation Accuracy (%)Day 1Proportion of correct laboratory/imaging interpretations or appropriate next-test selections, per AI tool and pooled; stratified by difficulty tier. Unit: percent (0-100).
AI Diagnostic Accuracy (%)Day 1Proportion of vignettes with a correct primary diagnosis produced by each anonymized AI tool and pooled across tools. Correctness is defined against a pre-specified reference answer key; results are also stratified by pre-defined difficulty tiers (easy/moderate/difficult/very difficult). Unit of measure: percent (0-100).
AI Medication-Dosing Accuracy (%)Day 1Proportion of dose recommendations meeting pediatric standards (weight- or BSA-based ranges, route, frequency) per reference rubric, per AI tool and pooled; stratified by difficulty tier. Unit: percent (0-100).

Secondary

MeasureTime frameDescription
AI Response Time (seconds per vignette)Day 1Time from vignette display to final AI output, reported per tool and pooled; also by difficulty tier. Unit: seconds.
Change in Physician Diagnostic Accuracy (percentage points) (Group 2 only)Day 1: Baseline (pre-AI) and immediate Post-AI within the same session (0-15 min after baseline).Post-AI accuracy minus pre-AI accuracy per participant on the same case set; also categorized as beneficial (incorrect→correct), harmful (correct→incorrect), or no change. Accuracy is the proportion of cases with a correct final diagnosis according to a prespecified answer key.
Net Benefit Index of AI Exposure (percentage points) (Group 2 only)Day 1Beneficial change rate (incorrect→correct) minus harmful change rate (correct→incorrect) for diagnostic items; sensitivity analyses for dosing/interpretation. Unit: percentage points.
Confidence Shift (Δ on a 1-10 scale) (Group 2 only)Day 1: Baseline (pre-AI) and immediate Post-AI within the same session (0-15 min after baseline).Post-AI self-rated confidence minus pre-AI confidence; association with correctness is examined. Unit: scale points (-9 to +9).
Answer-Change Frequency (%) (Group 2 only)Day 1Proportion of vignettes for which physicians revised their initial answer after AI suggestions; reported overall and by difficulty tier. Unit: percent (0-100).

Countries

Turkey (Türkiye)

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Sep 6, 2026