Skip to content

Generative AI-Assisted Clinical Decision Support for Medical Intensive Care Unit Physicians

Evaluation of the Feasibility and Effectiveness of Generative AI-Assisted Multidisciplinary Decision Support in Medical Intensive Care: A Pilot Randomized Controlled Trial

Status
Completed
Phases
Unknown
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT07708155
Enrollment
15
Registered
2026-07-16
Start date
2025-12-01
Completion date
2026-05-31
Last updated
2026-07-21

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Clinical Decision-making, Clinical Decision Support

Keywords

ChatGPT, Generative Artificial Intelligence, Large Language Model, Clinical Decision Support, Physician Decision-Making, Medical Intensive Care Unit, Critical Care, Usability, Cluster-Randomized Crossover Study, Acceptability

Brief summary

This pilot study evaluated the feasibility and usefulness of generative artificial intelligence (AI) as a clinical decision-support tool for physicians working in a medical intensive care unit. Participating physicians were assigned by work period to either use a generative AI system in addition to usual clinical information resources or to use usual resources without generative AI. The assigned condition was then switched so that participants experienced both approaches. During the AI-assisted periods, physicians used de-identified clinical information and considered the AI-generated responses as reference information. All final clinical decisions remained the responsibility of the treating physicians. The study assessed acceptability, usability, satisfaction, perceived decision support, workload, confidence, and learning experience through repeated questionnaires.

Detailed description

This was a single-center, open-label, pilot cluster-randomized crossover study involving physicians working in a medical intensive care unit. Each participating physician was observed during a scheduled one-month rotation in the medical intensive care unit. At the beginning of each monthly rotation, participating physicians were divided into two clusters. The clusters were randomized to begin with either the ChatGPT-assisted condition or the control condition. After approximately two weeks, each cluster crossed over to the alternate condition for the remainder of the one-month rotation. This design allowed participating physicians to experience both study conditions within the same rotation. During the AI-assisted condition, physicians were encouraged to use ChatGPT (OpenAI) as a reference tool to support clinical information review and decision-making. Only non-identifiable clinical information was permitted to be entered into ChatGPT. Patient names, medical record numbers, contact information, and other information that could directly identify an individual patient were not entered. Physicians summarized clinically relevant information in their own words and considered the responses generated by ChatGPT when planning patient management. The Situation-Background-Assessment-Recommendation framework was recommended as an optional structure for organizing clinical information, but its use was not mandatory. Physicians were otherwise free to formulate their queries and interact with ChatGPT according to their clinical needs. A suggested prompt encouraged ChatGPT to present multiple management options, together with their rationale, potential benefits and risks, relevant supporting evidence, and areas of uncertainty. ChatGPT did not make or implement clinical decisions. During the control condition, physicians used usual information resources, including discussions with other clinicians, multidisciplinary rounds, consultations, textbooks, clinical practice guidelines, PubMed, and other established clinical reference services, without using ChatGPT or other generative AI tools for study-related clinical decision support. All diagnostic and treatment decisions were made independently by the treating physicians. Repeated questionnaires assessed satisfaction, decision-making experience, confidence, perceived efficiency, workload, educational value, and other aspects of clinical decision support. At study completion, participants also evaluated usability, satisfaction, perceived learning, reliance on ChatGPT, intention for future use, and the extent to which ChatGPT-generated suggestions were reflected in their clinical plans.

Interventions

OTHERGenerative AI-Assisted Clinical Decision Support

During the assigned period, physicians were encouraged to use ChatGPT (OpenAI) as a generative AI-based reference tool to support clinical information review and decision-making.

OTHERUsual Clinical Information Resources

During the control period, physicians used usual clinical information resources, including discussions with other clinicians, multidisciplinary rounds, specialty consultations, textbooks, clinical practice guidelines, PubMed, and established clinical reference services. No generative AI tool was used for clinical decision support during this period.

Sponsors

Seoul National University Hospital
Lead SponsorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
CROSSOVER
Primary purpose
HEALTH_SERVICES_RESEARCH
Masking
NONE

Intervention model description

The unit of randomization was the physician cluster within each one-month medical intensive care unit rotation. At the beginning of each one-month medical intensive care unit rotation, participating physicians were divided into two clusters. The physician clusters were randomized to one of two intervention sequences: the ChatGPT-assisted condition followed by the control condition, or the control condition followed by the ChatGPT-assisted condition. Each period lasted approximately two weeks, after which the clusters crossed over to the alternate condition. Thus, all participating physicians experienced both study conditions during the same monthly rotation.

Eligibility

Sex/Gender
ALL
Age
19 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

* Age 19 years or older. * Physicians, including residents, fellows, and attending physicians, working in the medical intensive care unit at Seoul National University Hospital. * Scheduled to work as a primary treating physician for at least 5 days during a planned observation period. * Able and willing to provide written informed consent.

Exclusion criteria

* Did not provide written informed consent. * Withdrew consent from study participation.

Design outcomes

Primary

MeasureTime frameDescription
Daily Physician Satisfaction ScoreAt the end of each working day during the one-month medical intensive care unit rotationThe score of 10 questionnaire items assessing physicians' satisfaction with their daily clinical work, modified from Shore and Franks (1986) and Suchman et al. (1993). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). Negatively worded items were reverse-scored. The mean score ranges from -2 to +2, with higher scores indicating greater satisfaction.
Daily Clinical Decision-Making ScoreAt the end of each working day during the one-month medical intensive care unit rotationThe score of 6 questionnaire items assessing satisfaction with the clinical decision-making process, perceived decision difficulty, clarity of the preferred treatment, availability of relevant information, and identification of factors affecting the decision. Items were modified from Gedney (1994) and Dolan (1999) and rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). Negatively worded items were reverse-scored. The mean score ranges from -2 to +2, with higher scores indicating a more favorable decision-making experience.
Perceived Quality Score for ChatGPTAt the end of the one-month medical intensive care unit rotation, after completion of both crossover periodsThe score of 8 questionnaire items assessing the perceived information quality, system quality, and service quality of ChatGPT, modified from Pillong et al. (2025). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). The negatively worded response-time item was reverse-scored. The mean score ranges from -2 to +2, with higher scores indicating better perceived quality.
Generative AI Usability ScoreAt the end of the one-month medical intensive care unit rotation, after completion of both crossover periodsThe score of 3 questionnaire items assessing ease of use, ease of learning, and clarity of interaction with Generative AI (ChatGPT), modified from Pillong et al. (2025). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). The mean score ranges from -2 to +2, with higher scores indicating greater usability.
Satisfaction Score for Generative AI UseAt the end of the one-month medical intensive care unit rotation, after completion of both crossover periodsThe score of 6 questionnaire items assessing the perceived usefulness, productivity, effectiveness, overall satisfaction, appropriateness, and intention to reuse Generative AI (ChatGPT), modified from Pillong et al. (2025). Each item was rated on a 5-point Likert scale from -2 (strongly disagree) to +2 (strongly agree). The mean score ranges from -2 to +2, with higher scores indicating greater satisfaction.

Countries

South Korea

Contacts

PRINCIPAL_INVESTIGATORMinju Han, M.D.

Seoul National University Hospital

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Jul 22, 2026