Skip to content

Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models

Comparative Analysis of Diagnostic Accuracy and Case Difficulty Assessment of Three Large Language Models in Endodontics Using Expert Consensus as the Reference Standard

Status
Not yet recruiting
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07732985
Enrollment
342
Registered
2026-07-29
Start date
2026-09-01
Completion date
2027-01-01
Last updated
2026-07-29

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Pulp and Periapical Tissue Disease

Brief summary

This prospective, blinded diagnostic accuracy study aims to compare the performance of three large language models-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-in endodontic diagnosis and case difficulty assessment. The models will be evaluated against expert consensus as the reference standard using standardized clinical data and periapical radiographs. Diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus will be assessed to determine the potential of LLMs as clinical decision-support tools in endodontics.

Detailed description

This prospective, blinded diagnostic accuracy study aims to compare the diagnostic accuracy and endodontic case difficulty assessment of three LLMs-ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7-with expert consensus as the reference standard. Adult patients requiring primary endodontic treatment or retreatment will undergo routine clinical examination, pulp sensibility testing, and periapical radiographic assessment. Standardized clinical information and radiographs will be provided to each LLM and to four experienced endodontists independently. Expert consensus, defined as agreement among at least three of four experts, will serve as the reference standard. The primary outcomes are diagnostic accuracy, sensitivity, specificity, and agreement with expert consensus for pulpal and periapical diagnosis and endodontic case difficulty assessment according to the American Association of Endodontists (AAE) criteria. The findings will provide evidence regarding the reliability of LLMs as clinical decision-support tools in endodontics and their potential role in improving diagnostic consistency and treatment planning.

Interventions

DEVICEChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, Claude Opus 4.7

Large language model used to analyze standardized clinical information and periapical radiographs to provide pulpal and periapical diagnosis and endodontic case difficulty assessment according to AAE criteria

Sponsors

Cairo University
Lead SponsorOTHER

Study design

Observational model
OTHER
Time perspective
PROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
16 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

1. Age above 16 years old. 2. Systemically healthy patient (ASA I or II). 3. Requiring endodontic treatment or retreatment 4. Patient's acceptance to participate in the study

Exclusion criteria

1. Medically compromised patients. 2. Pregnant women. 3. Traumatic dental injuries 4. Low quality periapical radiograph

Design outcomes

Primary

MeasureTime frameDescription
Diagnostic accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 for pulpal and periapical diagnosisAt baselineDiagnostic performance of each large language model compared with the expert consensus reference standard. Accuracy, sensitivity, specificity, and agreement (Cohen's kappa) will be calculated according to the American Association of Endodontists (AAE) diagnostic criteria

Secondary

MeasureTime frameDescription
Accuracy of ChatGPT (GPT-5.5 Pro), Gemini 3.1 Pro, and Claude Opus 4.7 in endodontic case difficulty assessmentAt baselineAgreement between each large language model and the expert consensus in classifying endodontic case difficulty according to the AAE Endodontic Case Difficulty Assessment Guidelines. Overall accuracy and weighted Cohen's kappa will be calculated.

Contacts

CONTACTMariam Ahmed Hossam
mariam.ahmed.hosam@dentistry.cu.edu.eg00201110913251

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Jul 30, 2026