Skip to content

Potentials and Limitations of Large Language Models in the Processing of Routine Clinical Data in Spine Surgery

Potentials and Limitations of Large Language Models in the Processing of Routine Clinical Data in Spine Surgery

Status
Active, not recruiting
Phases
Unknown
Study type
Observational
Source
DRKS
Registry ID
DRKS00041009
Enrollment
200
Registered
2026-07-15
Start date
2026-07-15
Completion date
Unknown
Last updated
2026-08-03

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Routine clinical data from patients with degenerative or traumatic spinal disorders

Interventions

Group 1: This retrospective cohort study will analyze pseudonymized routine clinical data from patients who underwent spine surgery between 2020 and 2026. Operative reports, clinical documentation, ra

Sponsors

GFO Kliniken Rhein-Berg
Lead Sponsor

Eligibility

Sex/Gender
All
Age
18 Years to No maximum

Inclusion criteria

Inclusion criteria: - Patients who underwent spinal surgery, particularly for degenerative or traumatic spinal disorders. - Availability of a complete written operative report generated as part of routine clinical care. - Availability of sufficient clinical information, including: preoperative diagnosis and surgical indication, clinical examination findings, relevant comorbidities, preoperative clinical documentation, and preoperative imaging studies (e.g., CT, MRI, or radiographs) that were used for routine clinical decision-making.

Exclusion criteria

Exclusion criteria: - Missing or incomplete operative reports. - Insufficient clinical documentation preventing a valid assessment of the LLM-generated outputs (e.g., missing preoperative diagnosis, imaging studies, or clinical documentation). - Missing or insufficient clinical or radiological source data required for the predefined study objectives. - Records containing substantial documentation errors or inconsistencies that preclude reliable evaluation of the generated outputs. - Non-pseudonymized datasets or datasets that cannot be processed in accordance with applicable data protection regulations. - Documents not generated as part of routine clinical care (e.g., documents created exclusively for research purposes).

Design outcomes

Primary

MeasureTime frame
Primary Endpoint The primary endpoint is the factual accuracy of the texts generated or transformed by large language models (LLMs) compared with the underlying routine clinical data. The generated outputs will be evaluated independently by two board-certified spine surgeons using a standardized assessment instrument. The following criteria will be assessed: - Agreement between the generated content and the underlying clinical data - Completeness of relevant medical information (e.g., indication for surgery, key surgical steps, implants, and complications) - Absence of factual inaccuracies - Absence of hallucinations or information not supported by the source data Overall performance will be rated using a 5-point Likert scale: 1= Major factual errors or hallucinations; not clinically usable 2= Multiple relevant errors or omissions 3= Mostly accurate with some clinically relevant inaccuracies 4= Largely accurate with only minor deviations 5= Completely accurate with no relevant errors or hallucinations

Secondary

MeasureTime frame
Secondary Endpoints - Completeness of relevant medical information (e.g., surgical indication, key operative steps, and complications) - Clinical usability of the generated content as assessed by medical experts - Clarity and readability of the generated texts, particularly for patient-oriented or lay-language outputs - Quality of content organization (e.g., logical structure and overall clarity) - Clinical plausibility of the presented surgical treatment options - Transparency and comprehensibility of the underlying clinical reasoning (quality of the provided rationale) - Occurrence of potential errors or safety-related issues (e.g., unsafe recommendations or omission of relevant contraindications) - Comparison of the original clinical documents and LLM-generated outputs with respect to all of the above evaluation criteria

Countries

Germany

Contacts

Public ContactHabib Bendella

51465

habib.bendella@gfo-kliniken-rhein-berg.de02202 938-0

Outcome results

None listed

Source: DRKS (via WHO ICTRP) · Data processed: Aug 10, 2026