Skip to content

Multimodal Deep Learning Model for Multi-task Diagnosis and Triage Suggestions of Ophthalmic Diseases

Development and Validation of Multimodal Deep Learning Model for Autonomous Diagnosis, Generative Reporting, and Specialist Referral in Ophthalmic Diseases: An International Multicenter Cohort Study

Status
Recruiting
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07447973
Enrollment
2000
Registered
2026-03-04
Start date
2026-03-10
Completion date
2027-12-31
Last updated
2026-07-07

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Anterior Segment Diseases

Keywords

Anterior Segment Diseases, Artificial Intelligence

Brief summary

Accurate and comprehensive interpretation of anterior segment diseases from slit-lamp and smartphone photographs remains a clinical challenge due to the limited specificity and structure of existing Artificial Intelligence tools. The purpose of this international, multicenter clinical trial is to developed and validated an agent-based framework that integrates vision-language models and large language models to enhance the diagnostic workflow of anterior segment diseases.

Interventions

DIAGNOSTIC_TESTMultimodal Vision-language Model Diagnosis

Multimodal Vision-language Model for Multi-task Diagnosis and Triage Suggestions of Ophthalmic Diseases Patients presenting with complaints of anterior segment diseases first complete a slit-lamp examination or take a mobile phone eye photograph. A multimodal vision-language model uses patient-related images (such as selfies and eye exam photos) to make an intelligent diagnosis. The diagnosis is kept private. The patient then seeks medical attention and undergoes a clinical examination by an experienced clinician. A second experienced clinician then reviews the clinical diagnosis. If the diagnosis agrees, it is considered the gold standard. If there is a discrepancy in the diagnosis, the consensus between the two clinicians is used as the gold standard.

Sponsors

Guangdong Provincial People's Hospital
Lead SponsorOTHER

Study design

Observational model
OTHER
Time perspective
CROSS_SECTIONAL

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

1. Informed consent obtained; 2. Participants should be sufficiently able to read, write, and understand Chinese or English; 3. For normal participants: individuals should have no concerns related to their eyes. 4. For participants with eye-related chief complaints: individuals should have specific concerns or issues related to their eyes.

Exclusion criteria

1. Incomplete clinical data to support final diagnosis; 2. Patients who, in the opinion of the attending physician or clinical study staff, are too medically unstable to participate in the study safely.

Design outcomes

Primary

MeasureTime frameDescription
Diagnostic accuracy of multimodal vision-language model.from July 2025 to September 2025For each patient, the diagnoses generated by the multimodal vision-language model and the clinical diagnosis provided by skilled clinicians were documented and compared. Consistency between the two diagnoses indicates the program's precision in clinical practice.

Countries

China

Contacts

CONTACTHonghua Yu
yuhonghua@gdph.org.cn+8618688888422

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Jul 8, 2026