Skip to content

Performance of Large Language Models for Structured Recognition and Refractive Prediction

Head-to-Head Evaluation of ChatGPT 4o, GPT-5, and DeepSeek for Structured Extraction, Toric IOL Recommendation, and Refractive Prediction

Status
Recruiting
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07183891
Enrollment
100
Registered
2025-09-19
Start date
2025-08-01
Completion date
2035-12-31
Last updated
2025-09-19

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Cataract

Keywords

cataract, toric IOL, astigmatism

Brief summary

We conducted a single-center, retrospective observational study to evaluate large language models (ChatGPT 4o, GPT-5, DeepSeek) for automated interpretation of de-identified IOLMaster 700 reports provided as raster images. Models produced structured biometric extraction, toric IOL recommendation, and refractive predictions (sphere, cylinder, axis). Primary outcomes included parameter-level agreement and refractive error metrics; secondary outcomes included decision-support performance for toric IOL selection and agreement on ordered T-codes. No clinical intervention was performed.

Detailed description

This study compares three large language models accessed in their native configurations, without fine-tuning or external tools. For each examination, the original IOLMaster 700 report image was supplied without manual annotation or pre-processing. A standardized instruction required: (i) structured extraction of AL, ACD, LT, WTW, K1/K2 and axes, ΔK, TK1/TK2 and axes, and ΔTK; (ii) binary toric candidacy and T-code according to institutional ALCON mapping; and (iii) refractive recommendations (sphere, cylinder, implantation axis). Each model generated three independent outputs per case. De-identification and IRB oversight (waiver of consent) were implemented according to institutional policy. The unit of enrollment is participants (n=54), with outcomes analyzed per eye (162 eyes) and per model generation where applicable.

Interventions

None listed

Sponsors

Jin Yang
Lead SponsorOTHER

Study design

Observational model
CASE_ONLY
Time perspective
RETROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
No

Inclusion criteria

-postoperative corrected distance visual acuity (CDVA) of 0.10 logMAR or better -an absolute IOL rotational stability of less than 10∘ at the 1-month follow-up examination

Exclusion criteria

* incomplete biometric data on the examination report; * a history of previous ocular surgery or ocular trauma * the occurrence of intraoperative complications, such as an anterior capsular tear or posterior capsular rupture * the development of significant postoperative complications, including but not limited to severe intraocular infection or inadequate pupillary dilation.

Design outcomes

Primary

MeasureTime frameDescription
Refractive prediction error for sphereAt index examinationMean absolute error (MAE, diopters) of model-predicted sphere versus clinical reference
Cohen's kappa with 95% CIs between modelAt index examination (single time point)Cohen's kappa with 95% CIs between model outputs and clinician-validated reference for per-parameter

Secondary

MeasureTime frameDescription
Cylinder prediction errorAt index examinationMean absolute error (MAE, diopters) of model-predicted Cylinder
Axis prediction errorAt index examinationMean absolute error (MAE, diopters) of model-predicted Axis

Countries

China

Contacts

Primary ContactXuanqiao Lin
1532483480@qq.com+8615088920668

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Feb 4, 2026