Skip to content

Performances of Large Language Models in Kidney Allograft Diagnostics

Performances of Large Language Models in Kidney Allograft Diagnostics

Status
Completed
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07004660
Enrollment
240
Registered
2025-06-04
Start date
2004-03-01
Completion date
2023-12-31
Last updated
2025-06-04

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Kidney Transplantation

Keywords

Large language model, Banff classification, Diagnosis, Kidney transplantation, Biopsy

Brief summary

Kidney allograft rejection diagnosis relies on the complex Banff classification, but its application is limited by variability and workload. Our group previously built a scripted automation system, though it required major expert input. This study assesses whether modern LLMs can achieve similar diagnostic performance using Banff-based prompts, without extensive manual engineering.

Detailed description

Kidney allograft rejection remains a leading cause of allograft failure. Histological diagnosis relies on the Banff classification, a complex and evolving rule based framework. While successive Banff working groups refined the guidelines over time, daily interpretation is still hampered by inter and intra pathologist variability and growing demands on renal pathologists. This is why our group previously built a fully scripted Banff automation system. However, this system demanded years of expert curation and bespoke code before reaching acceptable accuracy. Whether modern LLMs, which show high capabilities to generate consistent and transparent reasoning at scale, can match expert pathologists without such resource intensive engineering remains unknown. The present study was therefore designed to benchmark state of the art LLMs against consensus diagnoses from senior renal pathologists on a representative series of kidney allograft biopsies, and to explore whether properly engineered prompts can translate Banff rules into reliable, reproducible diagnostic output.

Interventions

None listed

Sponsors

Paris Translational Research Center for Organ Transplantation
Lead SponsorOTHER

Study design

Observational model
COHORT
Time perspective
PROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
0 Years to 100 Years
Healthy volunteers
No

Inclusion criteria

* Kidney recipients

Exclusion criteria

* Combined transplant

Design outcomes

Primary

MeasureTime frameDescription
Biopsy diagnosisThe biopsy will be protocol biopsies performed at 3 months and 1 year post transplant, or for-cause biopsies performed at any time post transplant.The diagnosis of the biopsy will be based on the latest Banff classification (2022). The diagnoses will include i) biopsies with nonspecific lesions or clean (n=40), ii) biopsies with antibody-mediated rejection (AMR) (n=40) among which 14 had acute AMR, 13 chronic active AMR and 13 chronic inactive AMR, iii) biopsies with T cell-mediated rejection (TCMR) (n=40), among which 20 had acute TCMR and 20 had chronic active TCMR, iv) biopsies with borderline for acute TCMR (n=40), v) biopsies with mixed rejection (n=40), and vi) biopsies with microvascular inflammation (MVI) (n=40) among which 20 had probable AMR and 20 had MVI without DSA and without C4d.

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Feb 4, 2026