Routine clinical data from patients with degenerative or traumatic spinal disorders
Conditions
Interventions
Sponsors
Eligibility
Inclusion criteria
Inclusion criteria: - Patients who underwent spinal surgery, particularly for degenerative or traumatic spinal disorders. - Availability of a complete written operative report generated as part of routine clinical care. - Availability of sufficient clinical information, including: preoperative diagnosis and surgical indication, clinical examination findings, relevant comorbidities, preoperative clinical documentation, and preoperative imaging studies (e.g., CT, MRI, or radiographs) that were used for routine clinical decision-making.
Exclusion criteria
Exclusion criteria: - Missing or incomplete operative reports. - Insufficient clinical documentation preventing a valid assessment of the LLM-generated outputs (e.g., missing preoperative diagnosis, imaging studies, or clinical documentation). - Missing or insufficient clinical or radiological source data required for the predefined study objectives. - Records containing substantial documentation errors or inconsistencies that preclude reliable evaluation of the generated outputs. - Non-pseudonymized datasets or datasets that cannot be processed in accordance with applicable data protection regulations. - Documents not generated as part of routine clinical care (e.g., documents created exclusively for research purposes).
Design outcomes
Primary
| Measure | Time frame |
|---|---|
| Primary Endpoint The primary endpoint is the factual accuracy of the texts generated or transformed by large language models (LLMs) compared with the underlying routine clinical data. The generated outputs will be evaluated independently by two board-certified spine surgeons using a standardized assessment instrument. The following criteria will be assessed: - Agreement between the generated content and the underlying clinical data - Completeness of relevant medical information (e.g., indication for surgery, key surgical steps, implants, and complications) - Absence of factual inaccuracies - Absence of hallucinations or information not supported by the source data Overall performance will be rated using a 5-point Likert scale: 1= Major factual errors or hallucinations; not clinically usable 2= Multiple relevant errors or omissions 3= Mostly accurate with some clinically relevant inaccuracies 4= Largely accurate with only minor deviations 5= Completely accurate with no relevant errors or hallucinations | — |
Secondary
| Measure | Time frame |
|---|---|
| Secondary Endpoints - Completeness of relevant medical information (e.g., surgical indication, key operative steps, and complications) - Clinical usability of the generated content as assessed by medical experts - Clarity and readability of the generated texts, particularly for patient-oriented or lay-language outputs - Quality of content organization (e.g., logical structure and overall clarity) - Clinical plausibility of the presented surgical treatment options - Transparency and comprehensibility of the underlying clinical reasoning (quality of the provided rationale) - Occurrence of potential errors or safety-related issues (e.g., unsafe recommendations or omission of relevant contraindications) - Comparison of the original clinical documents and LLM-generated outputs with respect to all of the above evaluation criteria | — |
Countries
Germany
Contacts
51465