Large Language Models, Peritoneal Dialysis (PD)
Conditions
Brief summary
This study is a randomized controlled trial based on simulated clinical cases, aiming to establish a standardized evaluation system for PD physicians, to assess the differences in PD management quality between a workflow assisted by a PD-specialized large language model and physician-only decision-making, and to identify potential risks (such as generating obviously erroneous or even harmful recommendations). This simulated clinical case framework not only supports standardized and blinded evaluation, but also provides preliminary evidence for the model's effectiveness and safety before its deployment in real clinical settings, while avoiding direct impact on real patients.
Interventions
The peritoneal dialysis-specialized large language model (PD-LLM) used in this study was jointly developed by the Department of Nephrology at the First Affiliated Hospital of Sun Yat-sen University and Digital Health China (DHC).
Sponsors
Study design
Eligibility
Inclusion criteria
1. Internal medicine or nephrology standardized training residents, licensed residents, or attending physicians. 2. Independent experience in PD management ≤ 3 years. 3. Provided signed informed consent and agreed to comply with the trial procedures.
Exclusion criteria
1. Direct involvement in the development or training of the specialized PD large language model, or in the construction of the clinical scenarios/ scoring criteria used in this trial. 2. Participation in the pilot testing of all clinical scenarios used in this trial. 3. Inability or unwillingness to access the study platform or use online resources during the study period. 4. Experienced PD experts.
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Management Reasoning | Within the 120-minute test period of the first test. | Percent correct (range: 0 to 100) for each case. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Domain-Specific Scores | Within the 120-minute test period of the first test. | Percent correct (range: 0 to 100) for each case in each domain. |
| Severity of Potential Harm | Within the 120-minute test period of the first test. | The severity of potential harm will be classified as none, mild-to-moderate, or severe. |
| Time per Scenario Case | Within the 120-minute test period of the first test. | The time (in seconds) participants spend per case. |
| Self-Reported Confidence per Case | Within the 120-minute test period of the first test. | Scale 1-10. 10 represents being very confident. |
| Difference between the Phys group's scores when tested with PD-LLM assistance and when assisted only by traditional search methods | After 1-2 months of washout of first test | range from -100 to 100 |
Countries
China