Skip to content

Agreement Between Large Language Model-Generated Treatment Recommendations With Guideline-Based and Tumor Board Decisions in Gastrointestinal Cancer

Concordance of Large Language Model-Generated Treatment Recommendations With Multidisciplinary Tumor Board and Guideline-Based Decisions in Gastrointestinal Cancer: A Retrospective Cohort Study

Status
Completed
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07592338
Acronym
KITuKo
Enrollment
30
Registered
2026-05-18
Start date
2025-01-01
Completion date
2026-02-25
Last updated
2026-05-18

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Colorectal Cancer, Gastric Cancer (GC), Pancreatic Cancer

Keywords

Colorectal Neoplasms, Stomach Neoplasms, Pancreatic Neoplasms, Artificial Intelligence, Clinical Decision Support Systems

Brief summary

The goal of this observational study is to learn whether a computer program can suggest cancer treatments that match expert recommendations for people with gastrointestinal cancer (cancer of the pancreas, stomach, or colon and rectum). The main questions it aims to answer are: * Do the treatment suggestions from the computer program match current medical guidelines? * Do these suggestions match decisions made by a multidisciplinary tumor board (a team of cancer specialists)? Researchers will review existing medical records from people who have already been treated for these cancers. They will enter key clinical information into a computer program that uses artificial intelligence (AI). The program will generate treatment suggestions for each case. Researchers will then compare these suggestions with: * guideline-based treatment recommendations * decisions made by the tumor board This study will help researchers understand whether AI tools could support doctors in making cancer treatment decisions in the future.

Detailed description

Gastrointestinal cancers require complex treatment planning that often involves surgery, systemic therapy, and multidisciplinary coordination. Clinical decision-making is typically guided by evidence-based recommendations and discussed in multidisciplinary tumor boards. However, the increasing complexity of treatment strategies and guideline frameworks can make consistent and reproducible decision-making challenging in routine clinical practice. Recent advances in artificial intelligence have enabled the development of large language models (LLMs) that can process structured clinical information and generate text-based recommendations. These systems may offer a scalable approach to support clinical workflows, but their ability to produce reliable and clinically appropriate treatment suggestions in oncology remains uncertain. This study evaluates the performance of an LLM-based system in the context of gastrointestinal oncology using retrospectively collected clinical case data. Structured case summaries derived from routine clinical documentation are used as standardized input. The model generates treatment recommendations under controlled conditions, allowing systematic comparison with established clinical reference standards. The analysis focuses on the level of agreement between model-generated recommendations and established decision-making frameworks. In addition, the study explores how model performance varies across different clinical scenarios, including varying levels of disease complexity. Particular attention is given to situations in which recommendations differ, in order to better understand potential limitations of the model and identify patterns that may be clinically relevant. Furthermore, the study examines the consistency of model outputs when the same clinical information is processed multiple times. This provides insight into the stability and reproducibility of the system, which are important considerations for potential real-world use. The findings of this study are intended to inform the potential role of LLM-based tools as supportive systems in clinical decision-making. The study does not evaluate clinical outcomes or patient benefit, but instead focuses on agreement with established standards and expert-driven decisions as an initial step in assessing feasibility and safety.

Interventions

OTHERTreatment recommendation according to official German cancer guideline

Detailed treatment recommendation according to the official guideline of the Association of the Scientific Medical Societies in Germany (AWMF; Arbeitsgemeinschaft der Wissenschaftlichen Medizinischen Fachgesellschaften),

OTHERTreatment recommendation of a LLM

Structured clinical case summaries were analyzed by a GPT-4-class large language model to generate treatment recommendations.

OTHERTreatment recommendation of a multidisciplinary tumor board

Detailed treatment recommendation according to the case-specific postoperative tumor board review.

Sponsors

Medizinische Hochschule Brandenburg Theodor Fontane
Lead SponsorOTHER

Study design

Observational model
COHORT
Time perspective
RETROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
No

Inclusion criteria

* Histologically confirmed pancreatic, gastric, or colorectal adenocarcinoma * Treatment discussed in a multidisciplinary tumor board

Exclusion criteria

* Non-adenocarcinoma histology

Design outcomes

Primary

MeasureTime frameDescription
Concordance with guideline-based managementAt the time of multidisciplinary tumor board evaluation up to 4 weeks after surgeryAgreement between LLM-generated recommendations and AWMF guideline-supported treatment strategies

Secondary

MeasureTime frameDescription
Concordance with multidisciplinary tumor board decisionsAt the time of multidisciplinary tumor board evaluation up to 4 weeks after surgeryAgreement between LLM-generated recommendations and tumor board treatment strategies
Reproducibility of LLM recommendations across repeated runsAt the time of multidisciplinary tumor board evaluation up to 4 weeks after surgeryStructured clinical case vignettes were entered into ChatGPT using a standardized prompt template. To assess within-model reproducibility, each clinical vignette was analyzed in 3 independent model sessions performed on different days using identical clinical input.
Characterization of discordant recommendations (e.g., overtreatment, undertreatment)At the time of multidisciplinary tumor board evaluation up to 4 weeks after surgeryOvertreatment was defined as an LLM-generated recommendation exceeding the intensity of the reference recommendation. Undertreatment was defined as omission of a recommended treatment or recommendation of a less intensive strategy.

Countries

Germany

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: May 19, 2026