Artificial Intelligence (AI), Common Systemic Diseases, Diagnostic Imaging
Conditions
Keywords
Multimodal Large Model, Deep Learning, Radiology, Multicenter Study, Diagnostic Performance
Brief summary
Following model development and locking, the fixed model is evaluated in prospectively collected CT cohorts from two centers. The study is observational and does not affect clinical care. A subset of cases is used in a randomized crossover reader study.
Detailed description
After model locking, CT data are prospectively collected at two centers for observational validation without retraining or parameter adjustment. Model outputs do not influence patient management. A subset of eligible cases is selected for the randomized crossover reader study.
Interventions
Radiologists interpret the medical images independently without any assistance from the AI model to establish a baseline performance.
Radiologists interpret the same set of medical images with the assistance of the multimodal medical imaging large model to evaluate the improvement in diagnostic performance.
Sponsors
Study design
Eligibility
Inclusion criteria
* Patients who underwent CT examinations for common systemic diseases. * Imaging data must have confirmed clinical reference standards, expert consensus, or pathological diagnosis. * Availability of complete DICOM format images with standard acquisition protocols.
Exclusion criteria
* Poor image quality (e.g., severe motion or metal artifacts) that precludes definitive diagnosis. * Cases with incomplete clinical or pathological reference standards. * Corrupted image files or duplicate cases.
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Case-level Diagnostic Accuracy and Area Under the ROC Curve (AUC) | Up to 1 week per evaluation period | Evaluation of case-level diagnostic accuracy (defined as the proportion of diagnostic decisions matching the clinical ground-truth label) and discrimination performance (measured by AUC) to compare unaided radiologist performance versus AI-assisted performance. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Diagnostic Efficiency (Reading and Reporting Time) | Up to 1 week per evaluation period | Measurement of diagnostic efficiency recorded as the time (in seconds) taken by radiologists to complete the case review and generate findings, with and without AI assistance. |
| Inter-rater Agreement (Fleiss' Kappa) | Up to 1 week per evaluation period | Assessment of diagnostic consensus and inter-rater consistency among participating radiologists measured using Fleiss' kappa (κ). |
| Clinical Report Quality and Semantic Accuracy Score | Up to 1 week per evaluation period | Assessment of AI-generated draft report quality evaluated by senior experts on a 5-point Likert scale (focusing on semantic accuracy and clinical relevance) and automated metrics (GREEN and ROUGE-L). |
Countries
China
Contacts
The Third Affiliated Hospital of Southern Medical University