Skip to content

Performance of an OCR-Prompt-LLM Integrated Workflow for Extracting Multi-dimensional Clinical Data in Ischemic Heart Disease

Performance of an OCR-Prompt-LLM Integrated Workflow for Extracting Multi-dimensional Clinical Data in Ischemic Heart Disease

Status
Completed
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07499830
Acronym
OPAL-CAD
Enrollment
308
Registered
2026-03-30
Start date
2026-02-23
Completion date
2026-03-02
Last updated
2026-03-30

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Artificial Intelligence (AI), Coronary Artery Disease, Data Collection

Brief summary

This research aims to evaluate a comprehensive AI-driven workflow for both clinical data extraction and diagnostic classification in coronary artery disease (CAD). Leveraging OCR and Large Language Models (LLMs), the system is designed to extract ten key clinical parameters (such as LVEF and lab results) and provide diagnostic subtypes (UA, STEMI, NSTEMI, CCS) directly from unstructured inpatient records. A man-machine comparative trial will be conducted using a test set of 308 patients, where the performance of the LLM-based workflow will be benchmarked against the average diagnostic accuracy and processing time of seven clinical physicians. The findings will provide evidence for the feasibility of using LLMs to enhance clinical data structuring and diagnostic efficiency in cardiology.

Interventions

DEVICEOCR-Prompt-LLM Information Extraction Workflow

The intervention is an automated clinical data management system integrating Optical Character Recognition (OCR), optimized Prompt Engineering, and Large Language Models (LLMs). The workflow processes unstructured inpatient records to extract 10 key clinical indicators (e.g., LVEF, CAD subtypes, medications) and classifies the patient into specific coronary artery disease categories (UA, STEMI, NSTEMI, CCS)

Standard manual process where experienced clinical physicians collect and interpret patient information from medical records. This serves as the human benchmark for comparing diagnostic accuracy and operational efficiency.

Sponsors

China National Center for Cardiovascular Diseases
Lead SponsorOTHER_GOV

Study design

Observational model
COHORT
Time perspective
RETROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

1. Patients aged 18 years and older. 2. Clinical records of patients who were previously enrolled in the AIM-CHD (for the pilot/prompt optimization set) or SMART-CHD (for the internal validation cohort) studies. 3. Patients diagnosed with, or suspected of having, coronary artery disease (CAD), including subtypes: Unstable Angina (UA), STEMI, NSTEMI, and Chronic Coronary Syndrome (CCS).

Exclusion criteria

1. Clinical records with severe data fragmentation or missing more than 50% of the key clinical indicators. 2. Handwritten medical records or low-quality scans that are illegible for Optical Character Recognition (OCR) processing. 3. Duplicate records or records with conflicting "Gold Standard" labels that cannot be reconciled by the expert committee.

Design outcomes

Primary

MeasureTime frameDescription
Overall Diagnostic and Extraction Accuracy RateThrough study completion, an average of 3 months.To calculate the overall accuracy rate of the LLM-based workflow across 308 cases (including the pilot set, internal validation cohort, and external validation cohort) for 10 clinical indicators (e.g., LVEF, blood glucose, etc.) and 4 diagnostic subtypes of coronary artery disease. Accuracy is defined as the proportion of cases where the LLM's extraction or diagnostic results are perfectly consistent with the 'Gold Standard' established by human clinical experts.

Countries

China

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Mar 31, 2026