Skip to content

Does AI Make Clinicians More Appropriately Confident? A Randomized Study in Preterm Birth Prediction

Does AI Make Clinicians More Appropriately Confident? A Randomized Study in Preterm Birth Prediction

Status
Recruiting
Phases
Unknown
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT07402668
Enrollment
125
Registered
2026-02-11
Start date
2026-02-03
Completion date
2026-07-01
Last updated
2026-07-01

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Artificial Intelligence (AI) in Diagnosis, Preterm Birth

Keywords

Preterm birth, Premature birth, Diagnostic calibration, Diagnostic accuracy, Diagnostic confidence, Artificial intelligence

Brief summary

The goal of this randomized questionnaire-based study is to evaluate how different presentations of artificial intelligence (AI) decision support influence clinical judgment among medical doctors working in obstetrics and gynecology when assessing the risk of spontaneous preterm birth using clinical case vignettes with cervical ultrasound images. The study specifically compares two AI presentation formats: a binary classification (preterm vs term birth) and an individualized risk estimate of preterm birth. The main questions it aims to answer are: * Which AI presentation format leads to better alignment between clinicians' confidence and decision accuracy (diagnostic calibration)? * Do different AI presentation formats lead to helpful or harmful changes in clinical decisions? Participants will complete an online questionnaire in which they review clinical cases, make diagnostic and management decisions, rate their diagnostic confidence before and after seeing the AI output, and report their trust in the AI.

Interventions

BEHAVIORALAI prediction (binary)

AI decision support based on cervical ultrasound providing a binary classification (preterm birth before 37 weeks or term birth) in addition to standard clinical information.

BEHAVIORALAI risk estimate (%)

AI decision support based on cervical ultrasound providing an estimate of preterm birth risk (%) in addition to standard clinical information.

Sponsors

Rigshospitalet, Denmark
Lead SponsorOTHER
Technical University of Denmark
CollaboratorOTHER
The Foundation of 17.12.1981
CollaboratorOTHER
Department of Computer Science, University of Copenhagen, Denmark
CollaboratorUNKNOWN
Copenhagen Academy for Medical Education and Simulation
CollaboratorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
PARALLEL
Primary purpose
OTHER
Masking
SINGLE (Subject)

Masking description

Participants (clinicians) are blinded to randomized allocation and are unaware that different versions of the AI decision support are being compared. They view only the AI output presented within their assigned condition. No independent outcome assessors are involved. Outcomes are derived using pre-specified, objective scoring rules.

Eligibility

Sex/Gender
ALL
Healthy volunteers
Yes

Inclusion criteria

* Medical doctors currently working in or training within the field of obstetrics and gynecology. * Experience performing transvaginal cervical ultrasound examinations.

Exclusion criteria

\- No prior experience performing transvaginal cervical ultrasound examinations.

Design outcomes

Primary

MeasureTime frameDescription
Clinician diagnostic calibration (accuracy-confidence alignment) after AI exposure.Immediately after AI exposure during a single questionnaire session (approximately 20 minutes).Agreement between post-AI decision correctness (0/1) and post-AI confidence rating (0-10) will be quantified using the Brier score. Confidence will be rescaled to 0-1 and squared differences between confidence and correctness will be averaged across cases to produce a participant-level score. Lower scores indicate better diagnostic calibration. Results will be compared between randomized arms.

Secondary

MeasureTime frameDescription
Helpful switch rate and harmful switch rate.Baseline (pre-AI) and immediately after AI exposure during a single questionnaire session (approximately 20 minutes).Proportion of cases with helpful and harmful switches calculated for each participant and compared between study arms. Helpful switch = incorrect pre-AI decision changing to correct post-AI decision. Harmful switch = correct pre-AI decision changing to incorrect post-AI decision.
Change in decision accuracy, confidence, and diagnostic calibration from pre-AI to post-AI.Baseline (pre-AI) and immediately after AI exposure during a single questionnaire session (approximately 20 minutes).Within-participant change from pre-AI to post-AI in decision accuracy (proportion of correct decisions), confidence rating, and diagnostic calibration. Differences will be compared between randomized arms and stratified by AI correctness.
Association between self-rated trust in AI and behavioral reliance on AI.Immediately after AI exposure during a single questionnaire session (approximately 20 minutes).Self-rated trust in the AI output will be measured using a numeric rating scale (0-10) after AI exposure for each case. Behavioral reliance will be quantified as the proportion of post-AI decisions concordant with the AI output. The relationship between trust ratings and behavioral reliance, including concordance when the AI is correct and incorrect, will be evaluated at the participant level and compared between randomized arms.
Follow-up cervical ultrasound planning.Baseline (pre-AI) and immediately after AI exposure during a single questionnaire session (approximately 20 minutes).Proportion of cases in which clinicians plan an additional cervical ultrasound (yes/no), summarized per participant and compared pre-post AI and between randomized arms.

Countries

Denmark

Contacts

CONTACTEmilie Pi F Sejer, MD
emilie.pi.fogtmann.sejer.01@regionh.dk0045 28890690
STUDY_CHAIRMartin G Tolsgaard, MD, PhD, DMSc

Department of Obstetrics and Gynecology, Copenhagen University Hospital - Rigshospitalet, Copenhagen, Denmark

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Jul 2, 2026