Clinical Competence, History Taking, Medical Education
Conditions
Keywords
Large Language Model, Artificial Intelligence, Virtual Patient, Gastroenterology, Medical Education, History Taking, Simulation Training, Clinical Skills, Automated Assessment
Brief summary
This randomized controlled trial evaluated a large language model-based Intelligent Simulated Patient System (ISPS) for gastroenterology history-taking education. Ninety medical students were randomly assigned in a 1:1 ratio to 4 weeks of conventional clinical learning or ISPS-assisted training. The ISPS used teacher-defined structured case information to support simulated patient interactions and generated item-level formative feedback using predefined scoring rubrics. The implemented system did not use external electronic health record retrieval, retrieval-augmented generation, a vector database, or semantic-similarity threshold scoring. The primary outcomes were post-intervention medical history-taking performance and clinical diagnostic accuracy. History-taking performance was assessed in a standardized Objective Structured Clinical Examination by two senior physicians who were blinded to group allocation. Secondary outcomes included patient-centered communication competence and participants' acceptance of the ISPS.
Detailed description
Background and Objective The Intelligent Simulated Patient System (ISPS) was developed to provide repeatable history-taking practice, automated formative assessment, and structured feedback for medical students learning gastroenterology. The study evaluated the reliability of the automated history-taking scoring system and the educational effectiveness of ISPS-assisted training. ISPS Intervention Teachers configured structured disease templates, patient-specific case information, examination tasks, scoring items, and reference diagnoses within the system. Students conducted multiround history-taking interviews with large language model-driven simulated patients. The simulated patients were instructed to respond only according to the teacher-defined patient information. After each training encounter, the ISPS evaluated history-taking completeness item by item against a predefined 100-point Medical History-Taking Assessment (MHTA) rubric and generated formative feedback. Patient-centered communication was assessed separately in four domains: exploring the patient's ideas, exploring the patient's concerns, exploring the impact of illness, and expressing understanding and support. All case information and scoring criteria used in the evaluated version were configured by teachers and provided directly to the model as structured prompt context. The system did not use external electronic health record retrieval, retrieval-augmented generation, a separate vector database, knowledge-graph retrieval, instruction fine-tuning, SBERT-based semantic-similarity threshold scoring, or rule-tree coverage scoring. Automated Scoring Validation Before evaluation of the educational intervention, the automated history-taking scoring system was assessed using 50 complete interview records. The ISPS scored each record using the predefined MHTA rubric. Two senior clinical educators independently scored the same records using the identical rubric and were blinded to the ISPS-generated scores and to each other's ratings. Agreement between ISPS-generated scores and the mean expert scores was evaluated using Pearson correlation, the intra-class correlation coefficient, and Bland-Altman analysis. Randomized Controlled Trial Ninety eligible medical students who had completed the relevant theoretical and clinical instruction provided written informed consent and were randomly assigned in a 1:1 ratio to the intervention or control group. Participants in the control group received 4 weeks of conventional clinical learning, including teaching rounds, case-based discussions, paper medical records, bedside interviews, and verbal feedback from senior physicians. Participants in the intervention group independently practiced gastroenterology history taking using the ISPS. After each training encounter, the system provided a formative MHTA score and item-level feedback. These training scores were used for formative learning and were not used as the primary post-intervention outcome. Outcome Assessment The primary outcomes were post-intervention MHTA score and clinical diagnostic accuracy. History-taking performance was evaluated independently of the ISPS training platform in a standardized Objective Structured Clinical Examination. Each participant interviewed a trained human standardized patient, and two senior physicians who were blinded to group allocation independently rated the participant using the 100-point MHTA rubric. The mean of the two physician ratings was used as the final MHTA score. Clinical diagnostic accuracy was determined by whether the participant's first diagnosis matched the predefined reference diagnosis for the examination case. Secondary outcomes included patient-centered communication competence and satisfaction with and acceptance of the ISPS. Communication competence was assessed using a four-domain scale based on the Understanding the Patient's Perspective domain of the SEGUE framework. Each domain was scored from 1 to 5, producing a total score ranging from 4 to 20. Participants in the intervention group completed a Technology Acceptance Model questionnaire after the intervention. Registration Information Participant enrollment and follow-up had been completed before this study record was first submitted to ClinicalTrials.gov. The record was subsequently revised to ensure that the intervention description and outcome measures accurately reflected the study as conducted. Previous versions remain available in the ClinicalTrials.gov Record History.
Interventions
Participants received conventional clinical education for four weeks, including bedside teaching, case-based learning, paper-based medical record review, and faculty-guided clinical discussions according to the standard curriculum.
Participants received 4 weeks of self-directed gastroenterology history-taking training using the Intelligent Simulated Patient System (ISPS). Participants selected teacher-configured clinical cases and conducted multiround history-taking interviews with large language model-driven simulated patients. Simulated-patient responses were constrained by teacher-defined structured patient information. After each training encounter, the ISPS evaluated history-taking completeness item by item using a predefined 100-point Medical History-Taking Assessment rubric and provided individualized formative feedback. Patient-centered communication competence was assessed separately from the complete interview transcript using a four-domain scale. These automated scores and feedback were used solely for formative training and were not used as the primary post-intervention outcome of the randomized trial. In the system version evaluated in this study, case information and scoring criteria were configure
Sponsors
Study design
Masking description
Participants and instructors were aware of group assignment because of the nature of the educational intervention. Outcome assessors who evaluated the OSCE performance were blinded to group allocation. Statistical analyses were also performed by an independent statistician blinded to treatment assignment.
Intervention model description
Participants were randomly assigned in a 1:1 ratio to either the Intelligent Simulated Patient System (ISPS) group or the conventional teaching group for four weeks.
Eligibility
Inclusion criteria
* Undergraduate medical students at Zhejiang Chinese Medical University who had completed their clinical clerkship and the required coursework in Diagnostics and Internal Medicine (Gastroenterology). * Postgraduate students in Clinical Medicine at the First Affiliated Hospital of Zhejiang Chinese Medical University (Zhejiang Provincial Hospital of Chinese Medicine). * Able and willing to provide written informed consent. * Willing to complete the entire 4-week study protocol and all outcome assessments.
Exclusion criteria
* Individuals with extensive prior clinical experience in gastroenterology that could substantially influence baseline performance. * Individuals with severe visual, hearing, or other impairments that would prevent effective participation in the Intelligent Simulated Patient System training. * Individuals unable to complete the required training sessions or the final Objective Structured Clinical Examination (OSCE). * Individuals who withdrew informed consent or had incomplete outcome assessment.
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Diagnostic accuracy | Immediately after the 4-week intervention | Diagnostic accuracy was defined as the proportion of participants who correctly identified the primary diagnosis during the standardized OSCE. |
| History-taking completeness score | Immediately after the 4-week intervention | History-taking completeness was evaluated during a standardized Objective Structured Clinical Examination (OSCE) using a validated rubric with a total score ranging from 0 to 100. Higher scores indicate more complete history-taking performance. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Communication Competence | Immediately after the 4-week intervention | Communication competence was assessed from the complete standardized history-taking interview transcript using a four-domain scale based on the Understanding the Patient's Perspective domain of the SEGUE framework. The four domains were exploring the patient's ideas, exploring the patient's concerns, exploring the impact of illness, and expressing understanding and support. Each domain was scored from 1 to 5, producing a total score ranging from 4 to 20. Higher scores indicated better patient-centered communication competence. |
| Acceptance of the Intelligent Simulated Patient System | Immediately after completion of the intervention | Participants in the intervention group completed a 19-item questionnaire based on the Technology Acceptance Model. Items were rated using a 5-point Likert scale and assessed perceived usefulness, perceived ease of use, perceived enjoyment, attitude toward use and behavioral intention, and perceived risks and limitations. Item scores were transformed to a scale ranging from 20 to 100. Higher scores indicated more favorable evaluations, except in the Risks and Limitations domain, in which higher scores indicated greater perceived concerns. |
Countries
China