Clinical Decision Support, Emergency Care, Triage
Conditions
Keywords
Emergency Care, Medical AI, Large Language Models, Triage, Anonymization, De-Identification, Clinical Decision Support, AI Real World Data Validation
Brief summary
This retrospective, non-interventional study evaluates the prognostic performance of open-weight Large Language Models (LLMs) in the setting of a German academic emergency department. Using a full census of all consecutive emergency department cases at University Hospital Cologne between 01 January 2023 and 31 December 2025 (approximately 100,000 cases), the study assesses whether LLMs can make reliable prognostic predictions (e.g., hospital admission, imaging, diagnosis, placement) based on the initial history, vital signs, and triage category. In addition, it quantifies how strongly automated anonymization and perturbation procedures affect the models' diagnostic accuracy. This is an Investigator-Initiated Trial (IIT) with no intervention on patients.
Detailed description
The study analyzes a retrospective cohort of all emergency department cases at the Central Emergency Department of University Hospital Cologne (01 January 2023 - 31 December 2025). Data originate from the hospital information system (HIS) and are provided in pseudonymized form via the Medical Data Integration Center (MeDIC) of University Hospital Cologne, acting as an independent trusted third party. Extracted data include sociodemographic data (age/year of birth, sex), clinical vital signs (blood pressure, heart rate, respiratory rate, oxygen saturation, temperature, GCS), medical free text (triage records, physician history and admission findings), and process/outcome data serving as the gold standard (ICD-10 diagnoses, imaging performed, admission status, timestamp/length of stay). LLM processing takes place on premise on local compute clusters or in a contractually secured enterprise environment with zero data retention; open-weight models are used. Two arms are compared: Arm A (original data) vs. Arm B (anonymized/perturbed/synthesized data). The primary endpoint is diagnostic accuracy (AUROC, F1 score) against the documented clinical outcome. Working hypotheses: (1) modern LLMs are non-inferior to the human assessment (non-inferiority); (2) modern anonymization procedures reduce model performance by less than 5% (relative performance loss).
Interventions
None listed
Sponsors
Study design
Eligibility
Inclusion criteria
* All consecutive treatment cases at the Central Emergency Department of University Hospital Cologne during the period 01 January 2023 - 31 December 2025 (full census, consecutive inclusion).
Exclusion criteria
* Documented objection to the scientific use of the data pursuant to Art. 21 General data protection Regulation (GDPR). * Cases lacking the minimum data required for analysis (triage/history and documented outcome).
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Diagnostic accuracy of the LLM predictions (AUROC, F1 score) compared with the clinical gold standard | From enrollment to the end of retrospective observation period at 1 year | assessment at the level of the individual emergency department encounter).\] |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Relative performance loss of the models between original data (Arm A) and anonymized/perturbed data (Arm B); hypothesis < 5%. | From enrollment to the end of retrospective observation period at 1 year | Relative loss of diagnostic accuracy (AUROC, F1 score) when the models are applied to anonymized/perturbed data (Arm B) compared with original data (Arm A), expressed as the relative percentage change. Non-inferiority is assumed if the relative performance loss is below 5%.\] |
| Sensitivity, specificity, positive predictive value(PPV)/negative predictive value (NPV) for binary endpoints and agreement of the triage assessment (Cohen's kappa / Krippendorff's alpha). | From enrollment to the end of retrospective observation period at 1 year | Agreement between LLM output and the documented reference for binary endpoints (e.g., admission yes/no), reported as sensitivity, specificity, and positive/negative predictive value; agreement on the ordinal triage category is reported using Cohen's kappa or Krippendorff's alpha |
Countries
Germany
Contacts
Department of Internal Medicine II, University Hospital Cologne