Artificial Intelligence (AI), Burnout, Documentation, Efficiency of Physician Documentation, Empathy, Quality Improvement, Therapeutic Alliance
Conditions
Keywords
Artificial Intelligence, Psychiatry, Ambient Scribing
Brief summary
The goal of this clinical trial is to learn if an artificial intelligence (AI) tool that listens to appointments and drafts clinical notes can help clinicians in psychiatric care. Researchers also want to know if the tool changes how patients experience their care. The study includes mental health care providers (physicians, psychologists/psychotherapists, nurses, and social workers) at Psychiatrie St. Gallen and the adult patients they see in their appointments. The main questions it aims to answer are: * Does the AI tool lower the time clinicians spend writing notes? * Are notes written with the AI tool as good as or better than notes written without it? * Do patients notice a difference in their relationship with their clinician when the tool is used? Researchers will compare clinicians who use the AI tool to clinicians who keep writing notes their usual way, before and during the intervention period, to see if writing time, note quality, and the patient-rated therapeutic relationship differ between the two groups. Clinicians will: Join one of two groups. One group starts using the AI tool on August 14, 2026. The other group keeps writing notes the usual way. Both groups will be observed until the end of November 2026, with note quality evaluation completed in December 2026. Clinicians who cannot contribute to all measurements may, with their consent, remain in the study for automated measurement of documentation time only. Patients will: * Be invited to fill out a short, anonymous survey about their experience with their clinician. * Be asked to agree, in writing, that researchers may review the report from their initial assessment appointment.
Interventions
Listener Tool of the SwissGPT application (AlpineAI AG). After the patient's verbal consent, the provider records the consultation audio via mobile phone or laptop. All audio and text data are processed exclusively on AlpineAI's own GPU infrastructure located in Switzerland and are encrypted with individual keys per customer and per data type. Transcription uses the speech-to-text models Whisper 3 Turbo; from the transcript, a draft clinical note is generated by open-weight large language model Mistral Large 3 based on a standardized prompt developed for the institution. The provider transfers the generated draft note unedited into the electronic health record (ORBIS) and performs all review and editing directly in the EHR, so that the effort of correcting AI-generated drafts is captured by the automated documentation-time measurement. Audio recordings and transcripts are automatically deleted after 30 days.
Sponsors
Study design
Masking description
Providers and patients cannot be masked. Outcome assessors for report quality (PDQI-9 specialist raters) are blinded to provider identity, group allocation, time point, and site through removal of all directly identifying information from the reports before rating. Because AI-drafted notes may exhibit recognizable structure, blinding to scribe use cannot be assumed to be complete; as a prespecified manipulation check, raters record a forced-choice guess of scribe use for every report, and guess accuracy will be reported.
Intervention model description
Nonrandomized, open-label controlled trial with parallel assignment at the level of mental health care providers, analyzed within a difference-in-differences (DiD) framework. Outcomes are measured in both groups before and during the intervention period. Providers group assignment is fixed at enrollment.
Eligibility
Inclusion criteria
The observed study population comprises mental health care providers (physicians, psychologists/psychotherapists, nurses, and social workers) employed at Psychiatrie St. Gallen. Provider Inclusion Criteria: * Readiness to use an artificial intelligence ambient scribe during patient consultations (intervention group) * Regular clinical consultations with adult psychiatric patients Two outcome-specific criteria apply: contribution to the report quality outcomes requires that the provider regularly conducts new-patient intakes and writes initial psychiatric assessment reports; contribution to the patient-reported outcomes requires working in the outpatient setting. Providers who cannot or no longer wish to contribute to any actively collected outcome may, with their consent, remain enrolled for passive documentation-time measurement only; for analysis, such providers stay in the group which was fixed at enrollment.. Provider
Exclusion criteria
* notice of termination of the employment contract given during the data collection period * Working in the dementia inpatient clinic Patients are not the observed population, but anonymous ratings of their subjective consultation experience are collected from them, and written consent for the review of initial assessment reports is obtained. Patient Inclusion Criteria: \- All adult patients visiting the outpatient clinic during the data collection periods for a consultation with a participating provider Patient
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Time spent on psychiatric documentation | 24 Weeks | Mean minutes spent writing clinical notes (progress notes, initial psychiatric assessment reports, and discharge letters), extracted automatically from EHR (ORBIS) audit logs as form-level editing durations, compared pre-intervention vs. during intervention between groups. |
| Documentation quality of initial psychiatric asessment reports | 24 Weeks | Quality of initial psychiatric assessment reports rated with the Physician Documentation Quality Instrument, 9-item version (PDQI-9). Each of the 9 items is scored on a 5-point Likert scale (1 = strongly disagree to 5 = strongly agree) and summed to a total score ranging from 9 (minimum) to 45 (maximum); higher scores indicate better documentation quality. Each report is independently rated by 2 of 8 experienced mental health specialists in a balanced rotating-pair design, blinded to provider identity, group, time point, and site. |
| Working Alliance | 5 weeks pre intervention and 5 weeks during the intervention. | Patient-rated working alliance measured with the Working Alliance Inventory-Short Revised, German patient version (WAI-SR). The 12 items are each scored on a 5-point scale (1 to 5) and summed to a total score ranging from 12 (minimum) to 60 (maximum); higher scores indicate a stronger (better) working alliance. Reported anonymously by patients after outpatient consultations. In addition to the two-sided superiority test, non-inferiority is prespecified: a lower bound of the two-sided 95% CI of the difference-in-differences estimate above -3.5 WAI-SR sum-scale points will be interpreted as evidence against a clinically relevant deterioration of the alliance. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Time passed from initial psychiatric assessment note creation to its completion | 24 Weeks | Calendar time from the administrative creation of the initial psychiatric assessment note in the EHR to its completion, capturing documentation backlog. This outcome is separate from total active time spent writing the note. |
| Guideline-relevant information coverage and density in initial psychiatric assessment reports | 24 Weeks | Coverage and density of guideline-relevant psychiatric assessment content, scored by a large language model (SwissGPT) against a 71-item checklist derived from the American Psychiatric Association practice guidelines for the psychiatric evaluation of adults. Each item is scored as covered (1), partially covered (0.5), or not covered (0); report length in characters is used to compute information density. |
| Patient-perceived empathy of mental health care providers during consultations | 5 weeks pre intervention and 5 weeks during the intervention. | Patient-perceived empathy measured with the Consultation and Relational Empathy (CARE) measure. The 10 items are each scored on a 5-point scale (1 = poor to 5 = excellent) and summed to a total score ranging from 10 (minimum) to 50 (maximum); higher scores indicate greater perceived empathy (better). Completed anonymously by patients alongside the WAI-SR. A study-specific German translation is used; its psychometric properties will be examined and reported. |
| Documentation quality of ambient scribe-generated initial psychiatric assessment reports (raw AI output) | 15 Weeks | AI-generated initial psychiatric assessment reports, exactly as produced by the ambient scribe and before any human correction, rated by the generating providers (intervention group only) with the Provider Documentation Summarization Quality Instrument (PDSQI-9), an LLM-adapted 9-attribute instrument screening for hallucinations, omissions, redundancy, and biased language. Each attribute is scored on a 5-point scale (1 to 5) and summed to a total score ranging from 9 (minimum) to 45 (maximum); higher scores indicate better quality of the AI-generated summary. |
| Clinician Burnout | 5 weeks in the pre-intervention period and the last Week (Week 15) of the post-intervention period. | Work-related burnout of participating providers in both groups, measured with the German version of the Burnout Assessment Tool (BAT). The 23 core-symptom items (subscales: exhaustion, mental distance, cognitive impairment, emotional impairment) are each scored on a 5-point frequency scale (1 = never to 5 = always) and averaged to a core score ranging from 1.00 (minimum) to 5.00 (maximum); higher scores indicate more burnout symptoms (worse outcome). The 10 secondary-symptom items are scored and averaged on the same 1.00-5.00 metric, with higher scores likewise indicating a worse outcome. |
Countries
Switzerland
Contacts
Universität St. Gallen