Skip to content

AI Ambient Scribe in Psychiatry

Evaluating the Use of an Artificial Intelligence Ambient Scribe in a Swiss Psychiatric Clinic: Efficiency, Documentation Quality, and Therapeutic Alliance

Status
Active, not recruiting
Phases
Unknown
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT07777224
Acronym
AASP
Enrollment
47
Registered
2026-08-20
Start date
2026-07-13
Completion date
2026-12-31
Last updated
2026-08-20

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Artificial Intelligence (AI), Burnout, Documentation, Efficiency of Physician Documentation, Empathy, Quality Improvement, Therapeutic Alliance

Keywords

Artificial Intelligence, Psychiatry, Ambient Scribing

Brief summary

The goal of this clinical trial is to learn if an artificial intelligence (AI) tool that listens to appointments and drafts clinical notes can help clinicians in psychiatric care. Researchers also want to know if the tool changes how patients experience their care. The study includes mental health care providers (physicians, psychologists/psychotherapists, nurses, and social workers) at Psychiatrie St. Gallen and the adult patients they see in their appointments. The main questions it aims to answer are: * Does the AI tool lower the time clinicians spend writing notes? * Are notes written with the AI tool as good as or better than notes written without it? * Do patients notice a difference in their relationship with their clinician when the tool is used? Researchers will compare clinicians who use the AI tool to clinicians who keep writing notes their usual way, before and during the intervention period, to see if writing time, note quality, and the patient-rated therapeutic relationship differ between the two groups. Clinicians will: Join one of two groups. One group starts using the AI tool on August 14, 2026. The other group keeps writing notes the usual way. Both groups will be observed until the end of November 2026, with note quality evaluation completed in December 2026. Clinicians who cannot contribute to all measurements may, with their consent, remain in the study for automated measurement of documentation time only. Patients will: * Be invited to fill out a short, anonymous survey about their experience with their clinician. * Be asked to agree, in writing, that researchers may review the report from their initial assessment appointment.

Interventions

DEVICEArtificial Intelligence Ambient Scribe

Listener Tool of the SwissGPT application (AlpineAI AG). After the patient's verbal consent, the provider records the consultation audio via mobile phone or laptop. All audio and text data are processed exclusively on AlpineAI's own GPU infrastructure located in Switzerland and are encrypted with individual keys per customer and per data type. Transcription uses the speech-to-text models Whisper 3 Turbo; from the transcript, a draft clinical note is generated by open-weight large language model Mistral Large 3 based on a standardized prompt developed for the institution. The provider transfers the generated draft note unedited into the electronic health record (ORBIS) and performs all review and editing directly in the EHR, so that the effort of correcting AI-generated drafts is captured by the automated documentation-time measurement. Audio recordings and transcripts are automatically deleted after 30 days.

Sponsors

University of St.Gallen
CollaboratorOTHER
Psychiatrie St. Gallen
Lead SponsorOTHER_GOV

Study design

Allocation
NON_RANDOMIZED
Intervention model
PARALLEL
Primary purpose
HEALTH_SERVICES_RESEARCH
Masking
SINGLE (Outcomes Assessor)

Masking description

Providers and patients cannot be masked. Outcome assessors for report quality (PDQI-9 specialist raters) are blinded to provider identity, group allocation, time point, and site through removal of all directly identifying information from the reports before rating. Because AI-drafted notes may exhibit recognizable structure, blinding to scribe use cannot be assumed to be complete; as a prespecified manipulation check, raters record a forced-choice guess of scribe use for every report, and guess accuracy will be reported.

Intervention model description

Nonrandomized, open-label controlled trial with parallel assignment at the level of mental health care providers, analyzed within a difference-in-differences (DiD) framework. Outcomes are measured in both groups before and during the intervention period. Providers group assignment is fixed at enrollment.

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

The observed study population comprises mental health care providers (physicians, psychologists/psychotherapists, nurses, and social workers) employed at Psychiatrie St. Gallen. Provider Inclusion Criteria: * Readiness to use an artificial intelligence ambient scribe during patient consultations (intervention group) * Regular clinical consultations with adult psychiatric patients Two outcome-specific criteria apply: contribution to the report quality outcomes requires that the provider regularly conducts new-patient intakes and writes initial psychiatric assessment reports; contribution to the patient-reported outcomes requires working in the outpatient setting. Providers who cannot or no longer wish to contribute to any actively collected outcome may, with their consent, remain enrolled for passive documentation-time measurement only; for analysis, such providers stay in the group which was fixed at enrollment.. Provider

Exclusion criteria

* notice of termination of the employment contract given during the data collection period * Working in the dementia inpatient clinic Patients are not the observed population, but anonymous ratings of their subjective consultation experience are collected from them, and written consent for the review of initial assessment reports is obtained. Patient Inclusion Criteria: \- All adult patients visiting the outpatient clinic during the data collection periods for a consultation with a participating provider Patient

Design outcomes

Primary

MeasureTime frameDescription
Time spent on psychiatric documentation24 WeeksMean minutes spent writing clinical notes (progress notes, initial psychiatric assessment reports, and discharge letters), extracted automatically from EHR (ORBIS) audit logs as form-level editing durations, compared pre-intervention vs. during intervention between groups.
Documentation quality of initial psychiatric asessment reports24 WeeksQuality of initial psychiatric assessment reports rated with the Physician Documentation Quality Instrument, 9-item version (PDQI-9). Each of the 9 items is scored on a 5-point Likert scale (1 = strongly disagree to 5 = strongly agree) and summed to a total score ranging from 9 (minimum) to 45 (maximum); higher scores indicate better documentation quality. Each report is independently rated by 2 of 8 experienced mental health specialists in a balanced rotating-pair design, blinded to provider identity, group, time point, and site.
Working Alliance5 weeks pre intervention and 5 weeks during the intervention.Patient-rated working alliance measured with the Working Alliance Inventory-Short Revised, German patient version (WAI-SR). The 12 items are each scored on a 5-point scale (1 to 5) and summed to a total score ranging from 12 (minimum) to 60 (maximum); higher scores indicate a stronger (better) working alliance. Reported anonymously by patients after outpatient consultations. In addition to the two-sided superiority test, non-inferiority is prespecified: a lower bound of the two-sided 95% CI of the difference-in-differences estimate above -3.5 WAI-SR sum-scale points will be interpreted as evidence against a clinically relevant deterioration of the alliance.

Secondary

MeasureTime frameDescription
Time passed from initial psychiatric assessment note creation to its completion24 WeeksCalendar time from the administrative creation of the initial psychiatric assessment note in the EHR to its completion, capturing documentation backlog. This outcome is separate from total active time spent writing the note.
Guideline-relevant information coverage and density in initial psychiatric assessment reports24 WeeksCoverage and density of guideline-relevant psychiatric assessment content, scored by a large language model (SwissGPT) against a 71-item checklist derived from the American Psychiatric Association practice guidelines for the psychiatric evaluation of adults. Each item is scored as covered (1), partially covered (0.5), or not covered (0); report length in characters is used to compute information density.
Patient-perceived empathy of mental health care providers during consultations5 weeks pre intervention and 5 weeks during the intervention.Patient-perceived empathy measured with the Consultation and Relational Empathy (CARE) measure. The 10 items are each scored on a 5-point scale (1 = poor to 5 = excellent) and summed to a total score ranging from 10 (minimum) to 50 (maximum); higher scores indicate greater perceived empathy (better). Completed anonymously by patients alongside the WAI-SR. A study-specific German translation is used; its psychometric properties will be examined and reported.
Documentation quality of ambient scribe-generated initial psychiatric assessment reports (raw AI output)15 WeeksAI-generated initial psychiatric assessment reports, exactly as produced by the ambient scribe and before any human correction, rated by the generating providers (intervention group only) with the Provider Documentation Summarization Quality Instrument (PDSQI-9), an LLM-adapted 9-attribute instrument screening for hallucinations, omissions, redundancy, and biased language. Each attribute is scored on a 5-point scale (1 to 5) and summed to a total score ranging from 9 (minimum) to 45 (maximum); higher scores indicate better quality of the AI-generated summary.
Clinician Burnout5 weeks in the pre-intervention period and the last Week (Week 15) of the post-intervention period.Work-related burnout of participating providers in both groups, measured with the German version of the Burnout Assessment Tool (BAT). The 23 core-symptom items (subscales: exhaustion, mental distance, cognitive impairment, emotional impairment) are each scored on a 5-point frequency scale (1 = never to 5 = always) and averaged to a core score ranging from 1.00 (minimum) to 5.00 (maximum); higher scores indicate more burnout symptoms (worse outcome). The 10 secondary-symptom items are scored and averaged on the same 1.00-5.00 metric, with higher scores likewise indicating a worse outcome.

Countries

Switzerland

Contacts

PRINCIPAL_INVESTIGATORMaja Dobelsek, MD

Universität St. Gallen

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Aug 21, 2026