Skip to content

A Randomized Controlled Trial of Ambient Artificial Intelligence Scribe Technologies

A Randomized Controlled Trial of Two Ambient Artificial Intelligence Scribe Technologies to Improve Documentation Efficiency and Reduce Physician Burnout

Status
Completed
Phases
Unknown
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT06792890
Acronym
AIScribe RCT
Enrollment
238
Registered
2025-01-27
Start date
2024-11-04
Completion date
2025-01-15
Last updated
2026-04-24

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Artificial Intelligence (AI), Physician Workflow

Keywords

Artificial Intelligence Scribe, Randomized Controlled Trial, Documentation Efficiency, Physician Burnout

Brief summary

This is a three-arm pragmatic RCT of 238 outpatient physicians at a large academic health system, randomized 1:1:1 to one of two AI scribe tools or a usual-care control group. The two-month study will observe and compare the effects of each tool prior to system-wide roll out of selected tool (anticipated Spring 2025). We will use covariate-constrained randomization to balance the arms in terms of physician baseline time in notes, survey-measured level of burnout, and clinic days per week. The primary purpose of the initiative is to improve quality, efficiency, and business operations at University of California, Los Angeles (UCLA) Health, and this initiative is not being done for research purposes. The results of this operational initiative will inform the widespread roll out of AI scribe tools across all providers within the UCLA Health System. Nevertheless, the UCLA study team plans to rigorously examine and publish the impact of this intervention across the health system, which is why the study team pre-registered the initiative.

Detailed description

This study will assess operational-oriented outcomes across all groups. Notably, all groups will eventually receive all interventions over time in this observational study of a randomized roll out of a QI initiative. Moreover, the primary purpose of this initiative is operational. In other words, based on the results of this initiative, one of these tools will be eventually selected and operationalized widely across the health system. Enrolled participants are randomized to one of three groups. Randomization was needed to overcome secular trends, seasonal and holiday effects in December, and other factors confounding the relationship between exposure to the AI tools and the outcomes. The primary aim of this study is to evaluate the impact of two ambient AI scribe technologies on clinician change from baseline time spent on EHR documentation, comparing each scribe to a control group. Secondary objectives include assessing the AI scribes' impact on clinician metrics such as burnout, physician satisfaction, and productivity. Additionally, the study team intends to perform an economic evaluation analysis of the tools to guide business decision making. The study team will also analyze physician reported effects of the AI tools on patient safety, equity, and any unintended consequences of the initiative.

Interventions

OTHERUse Nabla AI Scribe tool provided

AI Scribe technologies capture physician-patient conversations to create a transcript, then summarize the transcript in the form of a clinical notes. These tools are integrated into the EHR and automatically adds the generated text to the provider note. All physicians must inform patients about the recording and obtain their verbal consent, and instances of patients declining to consent are tracked. Nabla leverages its proprietary speech-to-text to transform the conversation into a written context, combined with HIPAA compliant Large Language Models (LLM) like Azure OpenAI's GPT-4. Nabla does not store any audio.

OTHERUse AI Scribe tool provided by Vendor B

AI Scribe technologies capture physician-patient conversations to create a transcript, then summarize the transcript in the form of a clinical notes. These tools are integrated into the EHR and automatically adds the generated text to the provider note. All physicians must inform patients about the recording and obtain their verbal consent, and instances of patients declining to consent are tracked.

Sponsors

University of California, Los Angeles
Lead SponsorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
SINGLE_GROUP
Primary purpose
HEALTH_SERVICES_RESEARCH
Masking
SINGLE (Subject)

Masking description

The outpatient physicians are not told which tool that they are assigned to.

Eligibility

Sex/Gender
ALL
Healthy volunteers
No

Inclusion criteria

* Ambulatory care physicians within the UCLA Health system who held at least one half-day of clinic per week

Exclusion criteria

* Trainee providers (e.g., residents, medical students) and allied healthcare professionals (e.g., RNs, PAs) * Attendings who work exclusively with trainees

Design outcomes

Primary

MeasureTime frameDescription
Change in the Time in Notes Per NoteStudy month 2The primary outcome measure is the change in provider mean time in notes per note in the second month of the trial from the providers baseline mean time in notes per note for the six months prior to enrollment. This change will be computed on the natural log scale. No patient level information will be collected for this outcome measure.

Secondary

MeasureTime frameDescription
Provider Burnout ScoreStudy month 2The Mini Z 2.0 Survey is a validated 10-item instrument designed to measure key factors influencing workplace satisfaction and burnout among healthcare professionals. Each item is scored on a Likert scale (1-5), with higher scores generally indicating more positive outcomes - greater job satisfaction, sufficiency of time for electronic medical record documentation, and lower levels of stress. For negatively framed items (e.g., stress due to the job or frustration with the electronic medical record), higher scores indicate lower levels of dissatisfaction. The total score ranges from 10 to 50, with scores ≥40 representing a joyful workplace. No patient level information will be collected for this outcome measure.
Provider Task Load ScoreStudy month 2Provider task load adapted from the NASA Task Load Index (TLX), a validated tool for assessing perceived workload across six sub-scales: mental demand, physical demand, temporal demand, performance, effort, and frustration. For this study, we adapted the TLX to focus on note-writing workload, including four sub-scales (mental demand, temporal demand, physical demand, and effort) as done previously. Each sub-scale is rated from 0 (low task load) to 100 (high task load) and summed together for a total score scale of 0 (low task load) to 400 (high task load), lower is better. No patient level information will be collected for this outcome measure.
Provider Professional FulfillmentStudy month 2The Professional Fulfillment Index (PFI) is a validated 16-item instrument that uses a 5-point Likert scale (0-4) to measure professional fulfillment, work exhaustion, and interpersonal disengagement. For this study, we utilize the 4-item work exhaustion subscale which is a mean of the 4-items within that subscale, where a low score (0) indicates a lower level of exhaustion and a high (4) score indicates greater level of exhaustion. No patient level information will be collected for this outcome measure.
Number of Physicians Who Are Considered Detractors, Passive, or PromotersStudy month 2Self-reported satisfaction survey that asks physicians to consider note accuracy, patient safety, equity, and other potential unintended consequences and rate their overall likelihood to recommend use of the tool on a 1-10 scale. Higher scores (10) indicate greater satisfaction and likelihood to recommend, whereas lower scores (1) indicate dissatisfaction and unlikelihood to recommend. Providers are grouped as "Promoters" if they respond 9-10, "Passive" if they respond 7-8, and "Detractors" if they respond with a value less than or equal to 6. This grouping matches commonly accepted "Net Promoter Score" groupings. No patient level information will be collected for this outcome measure.
Change in Provider RVUStudy month 2The study team will use physician-level billing information via RVU to determine their change in productivity from a retrospective baseline 6 months prior to enrollment. No patient level information will be collected for this outcome measure.
Change in EHR Signal (Activity) Data - Pajama TimeStudy month 2We will examine change from a retrospective baseline 6 months prior to enrollment in Signal metrics including pajama time per scheduled day. Using this data will determine how a providers time is utilized in the EHR. No patient level information will be collected for this outcome measure.
Change in EHR Signal (Activity) Data - Time Outside Scheduled HoursStudy month 2We will examine change from a retrospective baseline 6 months prior to enrollment in Signal metrics including time outside scheduled hours per scheduled day. Using this data will determine how a providers time is utilized in the EHR. No patient level information will be collected for this outcome measure.
Change in EHR Signal (Activity) Data - Time on Unscheduled DaysStudy month 2We will examine change from a retrospective baseline 6 months prior to enrollment in Signal metrics including time spent in the system on unscheduled days where . Using this data will determine how a providers time is utilized in the EHR. No patient level information will be collected for this outcome measure.

Countries

United States

Participant flow

Pre-assignment details

Of 331 providers expressing interest, 93 were excluded due to ineligibility, declining participation, incomplete surveys, or decision-making roles. The remaining 238 were randomized to the arms.

Baseline characteristics

Characteristic
Age, Customized
Age
>=65
4 Participants
Age, Customized
Age
Between 25 and 34
12 Participants
Age, Customized
Age
Between 35 and 44
39 Participants
Age, Customized
Age
Between 45 and 54
59 Participants
Age, Customized
Age
Between 55 and 64
17 Participants
Age, Customized
Age
Prefer not to answer
10 Participants
Clinic days per week3.45 days/week
STANDARD_DEVIATION 1.35
Ethnicity (NIH/OMB)
Hispanic or Latino
6 Participants
Ethnicity (NIH/OMB)
Not Hispanic or Latino
68 Participants
Ethnicity (NIH/OMB)
Unknown or Not Reported
8 Participants
Race/Ethnicity, Customized
Race
Asian
98 Participants
Race/Ethnicity, Customized
Race
Black
2 Participants
Race/Ethnicity, Customized
Race
Multiple/Other
6 Participants
Race/Ethnicity, Customized
Race
Prefer not to answer
28 Participants
Race/Ethnicity, Customized
Race
White
84 Participants
Region of Enrollment
United States
79 Participants
Sex/Gender, Customized
Sex
Female
144 Participants
Sex/Gender, Customized
Sex
Male
87 Participants
Sex/Gender, Customized
Sex
Prefer not to answer
2 Participants
Single-item burnout3.49 units on a scale
STANDARD_DEVIATION 0.81
Specialty
Medical Specialty
33 Participants
Specialty
Primary Care
101 Participants
Specialty
Surgical Specialty
17 Participants
Time-in-note4.78 minutes per note

Adverse events

Event typeEG000
affected / at risk
EG001
affected / at risk
EG002
affected / at risk
deaths
Total, all-cause mortality
0 / 790 / 790 / 80
other
Total, other adverse events
0 / 790 / 790 / 80
serious
Total, serious adverse events
0 / 790 / 790 / 80

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Apr 25, 2026