Artificial Intelligence (AI), Coronary Artery Disease, Data Collection
Conditions
Brief summary
This research aims to evaluate a comprehensive AI-driven workflow for both clinical data extraction and diagnostic classification in coronary artery disease (CAD). Leveraging OCR and Large Language Models (LLMs), the system is designed to extract ten key clinical parameters (such as LVEF and lab results) and provide diagnostic subtypes (UA, STEMI, NSTEMI, CCS) directly from unstructured inpatient records. A man-machine comparative trial will be conducted using a test set of 308 patients, where the performance of the LLM-based workflow will be benchmarked against the average diagnostic accuracy and processing time of seven clinical physicians. The findings will provide evidence for the feasibility of using LLMs to enhance clinical data structuring and diagnostic efficiency in cardiology.
Interventions
The intervention is an automated clinical data management system integrating Optical Character Recognition (OCR), optimized Prompt Engineering, and Large Language Models (LLMs). The workflow processes unstructured inpatient records to extract 10 key clinical indicators (e.g., LVEF, CAD subtypes, medications) and classifies the patient into specific coronary artery disease categories (UA, STEMI, NSTEMI, CCS)
Standard manual process where experienced clinical physicians collect and interpret patient information from medical records. This serves as the human benchmark for comparing diagnostic accuracy and operational efficiency.
Sponsors
Study design
Eligibility
Inclusion criteria
1. Patients aged 18 years and older. 2. Clinical records of patients who were previously enrolled in the AIM-CHD (for the pilot/prompt optimization set) or SMART-CHD (for the internal validation cohort) studies. 3. Patients diagnosed with, or suspected of having, coronary artery disease (CAD), including subtypes: Unstable Angina (UA), STEMI, NSTEMI, and Chronic Coronary Syndrome (CCS).
Exclusion criteria
1. Clinical records with severe data fragmentation or missing more than 50% of the key clinical indicators. 2. Handwritten medical records or low-quality scans that are illegible for Optical Character Recognition (OCR) processing. 3. Duplicate records or records with conflicting "Gold Standard" labels that cannot be reconciled by the expert committee.
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Overall Diagnostic and Extraction Accuracy Rate | Through study completion, an average of 3 months. | To calculate the overall accuracy rate of the LLM-based workflow across 308 cases (including the pilot set, internal validation cohort, and external validation cohort) for 10 clinical indicators (e.g., LVEF, blood glucose, etc.) and 4 diagnostic subtypes of coronary artery disease. Accuracy is defined as the proportion of cases where the LLM's extraction or diagnostic results are perfectly consistent with the 'Gold Standard' established by human clinical experts. |
Countries
China