Skip to content

Improving the Reliability of LLMs as Medical Assistants for the General Public

Improving the Reliability of LLMs as Medical Assistants for the General Public: a Proof of Concept Simulation Trial

Status
Completed
Phases
Unknown
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT07651280
Acronym
LAMP-1
Enrollment
527
Registered
2026-06-16
Start date
2026-07-03
Completion date
2026-08-26
Last updated
2026-09-03

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Relevant Conditions Identification

Keywords

3M-6D education, Large Language Models, relevant conditions identification, Cognitive Load Theory

Brief summary

This study will evaluate whether three-minute six-dimensions education(3M-6D education) can improve the reliability of large language models as medical assistants for the general public. Participants will be randomly assigned to receive or not receive 3M-6D education and then use ChatGPT, Gemini, or non-AI information resources. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.

Detailed description

This randomized, controlled, proof-of-concept simulation trial will evaluate whether three-minute six-dimensions education (3M-6D education) can improve the reliability of large language models as medical assistants for the general public. Eligible participants will be randomly assigned in a 1:1:1:1:1 ratio to one of five study groups: the 3M-6D education GPT group, the GPT group, the 3M-6D education Gemini group, the Gemini group, or the control group. Participants in the 3M-6D education GPT and 3M-6D education Gemini groups will receive approximately three minutes of education before using ChatGPT or Gemini.Each participant will be randomly assigned one of 10 standardized clinical scenarios and complete a simulated counseling task in unrestricted natural language within approximately 10 minutes. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.

Interventions

BEHAVIORALthree minutes six dimensions education

3M-6D education is designed based on Cognitive Load Theory to reduce the cognitive burden on patients during medical interactions with AI and to improve the clarity and completeness of symptom reporting. Guided by cognitive load theory and the natural process physicians use to take medical histories, the investigators identified candidate information dimensions and developed a structured expression framework with six dimensions for public health queries through a Delphi expert consensus process. Participants were instructed to use the framework to describe their symptoms across these six dimensions; this process can typically be completed within three minutes, so the investigators call this approach three minutes six dimensions education (3M-6D education).

OTHERChatGPT

Participants use ChatGPT to complete a standardized simulated clinical scenarios in unrestricted natural language.

OTHERGemini

Participants use Gemini to complete a standardized simulated clinical scenarios in unrestricted natural language.

Sponsors

Capital Medical University
Lead SponsorOTHER
Xuanwu Hospital, Beijing
CollaboratorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
PARALLEL
Primary purpose
HEALTH_SERVICES_RESEARCH
Masking
SINGLE (Outcomes Assessor)

Masking description

Outcome assessors will be blinded to group assignment when evaluating participants' outcomes. Group information will be removed from the outcomes before assessment.

Intervention model description

Participants will be randomly assigned in a 1:1:1:1:1 ratio to one of five parallel groups: the 3M-6D education GPT group, the GPT group, the 3M-6D education Gemini group, the Gemini group, or the control group. Each participant will complete one standardized simulated clinical scenario.

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

1. Age 18 years or greater, male or female; 2. Completed primary school or higher education; 3. Able to use a smartphone or computer to complete online interaction; 4. No history of acute ischemic stroke, systemic lupus erythematosus, gastric ulcer, pneumonia, acute cardiac infarction, urinary tract infection, uterine fibroids, diabetes, osteoarthritis, or migraine. 5. Able to understand and comply with study procedures and to provide written informed consent.

Exclusion criteria

1. Currently or previously employed as a healthcare worker; 2. Previously received systematic medical training; 3. Currently involved in concurrent research that may interfere with the results of the present trial; 4. The investigator considered that the participant had other conditions that might affect compliance or preclude participation.

Design outcomes

Primary

MeasureTime frameDescription
Relevant conditions identification of the 3M-6D education GPT group compared with the GPT group1 hour.Relevant conditions identification is defined as the proportion of participants whose final response includes the expert-defined final diagnosis or a relevant differential diagnosis.
Disposition concordance of the 3M-6D education GPT group compared with the GPT group1 hour.Disposition concordance is defined as the proportion of participants whose final care recommendation matches the expert-defined level. The five levels are self-care, routine outpatient care, urgent outpatient care, emergency department visit, and emergency medical services.
Relevant conditions identification of the 3M-6D education Gemini group compared with the Gemini group1 hour.
Disposition concordance of the 3M-6D education Gemini group compared with the Gemini group1 hour.

Secondary

MeasureTime frameDescription
Relevant conditions identification of the 3M-6D education GPT group compared with the control group1 hour.
Relevant conditions identification of the 3M-6D education Gemini group compared with the control group1 hour.
Disposition concordance of the 3M-6D education GPT group compared with the control group1 hour.
Disposition concordance of the 3M-6D education Gemini group compared with the control group1 hour.
Red-flag identification in the 3M-6D education GPT group compared with the GPT group1 hour.Red-flag identification is defined as the proportion of participants whose final response includes the key warning signs that experts defined for the assigned scenario.
Red-flag identification in the 3M-6D education GPT group compared with the control group1 hour.
Red-flag identification in the 3M-6D education Gemini group compared with the Gemini group1 hour.
Red-flag identification in the 3M-6D education Gemini group compared with the control group1 hour.
NASA Task Load Index score of the 3M-6D education GPT group compared with the GPT group1 hour.NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
NASA Task Load Index score of the 3M-6D education GPT group compared with the control group1 hour.NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
NASA Task Load Index score of the 3M-6D education Gemini group compared with the Gemini group1 hour.NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
NASA Task Load Index score of the 3M-6D education Gemini group compared with the control group1 hour.NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
Relevant conditions identification of the 3M-6D education GPT group compared with the 3M-6D education Gemini group1 hour.
Disposition concordance of the 3M-6D education GPT group compared with the 3M-6D education Gemini group1 hour.
Red-flag identification in the 3M-6D education GPT group compared with the 3M-6D education Gemini group1 hour.
NASA Task Load Index score of the 3M-6D education GPT group compared with the 3M-6D education Gemini group1 hour.NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.

Countries

China

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Sep 4, 2026