Skip to content

Artificial Intelligence(AI) in the Emergency Department (ED) in Cologne

Retrospective Validation of Large Language Models (LLM) for the Prognostic Assessment of Clinical Parameters in Emergency Department and Evaluation of the Impact of Automated Anonymisation Methods

Status
Active, not recruiting
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07725965
Acronym
KINA-CO
Enrollment
100000
Registered
2026-07-24
Start date
2026-07-01
Completion date
2027-12-31
Last updated
2026-07-28

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Clinical Decision Support, Emergency Care, Triage

Keywords

Emergency Care, Medical AI, Large Language Models, Triage, Anonymization, De-Identification, Clinical Decision Support, AI Real World Data Validation

Brief summary

This retrospective, non-interventional study evaluates the prognostic performance of open-weight Large Language Models (LLMs) in the setting of a German academic emergency department. Using a full census of all consecutive emergency department cases at University Hospital Cologne between 01 January 2023 and 31 December 2025 (approximately 100,000 cases), the study assesses whether LLMs can make reliable prognostic predictions (e.g., hospital admission, imaging, diagnosis, placement) based on the initial history, vital signs, and triage category. In addition, it quantifies how strongly automated anonymization and perturbation procedures affect the models' diagnostic accuracy. This is an Investigator-Initiated Trial (IIT) with no intervention on patients.

Detailed description

The study analyzes a retrospective cohort of all emergency department cases at the Central Emergency Department of University Hospital Cologne (01 January 2023 - 31 December 2025). Data originate from the hospital information system (HIS) and are provided in pseudonymized form via the Medical Data Integration Center (MeDIC) of University Hospital Cologne, acting as an independent trusted third party. Extracted data include sociodemographic data (age/year of birth, sex), clinical vital signs (blood pressure, heart rate, respiratory rate, oxygen saturation, temperature, GCS), medical free text (triage records, physician history and admission findings), and process/outcome data serving as the gold standard (ICD-10 diagnoses, imaging performed, admission status, timestamp/length of stay). LLM processing takes place on premise on local compute clusters or in a contractually secured enterprise environment with zero data retention; open-weight models are used. Two arms are compared: Arm A (original data) vs. Arm B (anonymized/perturbed/synthesized data). The primary endpoint is diagnostic accuracy (AUROC, F1 score) against the documented clinical outcome. Working hypotheses: (1) modern LLMs are non-inferior to the human assessment (non-inferiority); (2) modern anonymization procedures reduce model performance by less than 5% (relative performance loss).

Interventions

None listed

Sponsors

University of Cologne
Lead SponsorOTHER

Study design

Observational model
OTHER
Time perspective
RETROSPECTIVE

Eligibility

Sex/Gender
ALL
Healthy volunteers
No

Inclusion criteria

* All consecutive treatment cases at the Central Emergency Department of University Hospital Cologne during the period 01 January 2023 - 31 December 2025 (full census, consecutive inclusion).

Exclusion criteria

* Documented objection to the scientific use of the data pursuant to Art. 21 General data protection Regulation (GDPR). * Cases lacking the minimum data required for analysis (triage/history and documented outcome).

Design outcomes

Primary

MeasureTime frameDescription
Diagnostic accuracy of the LLM predictions (AUROC, F1 score) compared with the clinical gold standardFrom enrollment to the end of retrospective observation period at 1 yearassessment at the level of the individual emergency department encounter).\]

Secondary

MeasureTime frameDescription
Relative performance loss of the models between original data (Arm A) and anonymized/perturbed data (Arm B); hypothesis < 5%.From enrollment to the end of retrospective observation period at 1 yearRelative loss of diagnostic accuracy (AUROC, F1 score) when the models are applied to anonymized/perturbed data (Arm B) compared with original data (Arm A), expressed as the relative percentage change. Non-inferiority is assumed if the relative performance loss is below 5%.\]
Sensitivity, specificity, positive predictive value(PPV)/negative predictive value (NPV) for binary endpoints and agreement of the triage assessment (Cohen's kappa / Krippendorff's alpha).From enrollment to the end of retrospective observation period at 1 yearAgreement between LLM output and the documented reference for binary endpoints (e.g., admission yes/no), reported as sensitivity, specificity, and positive/negative predictive value; agreement on the ordinal triage category is reported using Cohen's kappa or Krippendorff's alpha

Countries

Germany

Contacts

PRINCIPAL_INVESTIGATORVolker Burst, Prof.

Department of Internal Medicine II, University Hospital Cologne

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Jul 29, 2026