Skip to content

A Privacy-Preserving OCR-LLM System for Coronary Syndrome Subtyping From Admission HPI: Multicenter Validation in China and the US

Development and Multicenter Validation of a Privacy-Preserving OCR-LLM Pipeline for Four-Subtype Coronary Syndrome Classification Using Admission HPI Across Heterogeneous EHR Systems

Status
Not yet recruiting
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07449429
Acronym
OCR-LLM-CHD
Enrollment
10
Registered
2026-03-04
Start date
2026-02-28
Completion date
2026-03-08
Last updated
2026-03-04

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Acute Coronary Syndromes, Coronary Artery Disease (CAD) (E.G., Angina, Myocardial Infarction, and Atherosclerotic Heart Disease (ASHD)), Non-ST-Segment Elevation Myocardial Infarction (NSTEMI), ST-segment Elevation Myocardial Infarction (STEMI)

Brief summary

This study develops and validates a privacy-preserving OCR-LLM pipeline that converts admission history of present illness (HPI) records into structured coronary syndrome subtypes (STEMI, NSTEMI, unstable angina, and chronic coronary syndrome). The system first extracts text from de-identified HPI images using locally deployed OCR, then applies large language models with a fixed diagnostic prompt to generate subtype classification and evidence. Performance is evaluated in an internal validation cohort and multiple external datasets covering heterogeneous EHR templates, emergency department cases, and an English dataset from MIMIC-IV. A clinician usability study assesses changes in diagnostic accuracy and time with and without tool assistance.

Interventions

DEVICEOCR-Prompt-LLM Information Extraction and Classification Workflow (OCR-Prompt-LLM)

An automated clinical data management workflow integrating Optical Character Recognition (OCR), optimized prompt engineering, and large language models (LLMs). The system processes unstructured inpatient/ED records (primarily admission history of present illness and related narrative text) to extract prespecified key clinical indicators (e.g., left ventricular ejection fraction, coronary syndrome subtype, medications) and to classify cases into prespecified coronary artery disease categories (e.g., unstable angina, STEMI, NSTEMI, chronic coronary syndrome). The workflow outputs structured fields and a classification result with supporting evidence excerpts.

Standard manual process in which experienced clinicians review patient medical records and extract the same prespecified clinical indicators and coronary artery disease categories using routine clinical judgment and documentation review. This manual abstraction serves as the human benchmark for comparing diagnostic accuracy, completeness, and operational efficiency against the automated OCR-Prompt-LLM workflow.

Sponsors

China National Center for Cardiovascular Diseases
Lead SponsorOTHER_GOV

Study design

Observational model
COHORT
Time perspective
RETROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum

Inclusion criteria

Hospital encounters with admission HPI documenting sym 2a4afaef-9dc1-47fc-874f-9dffaf7… evant to coronary syndrome subtyping. Cases with sufficient documentation to assign one of four target subtypes (STEMI, NSTEMI, UA, CCS) by adjudication. \-

Exclusion criteria

Unclear subtype or incomplete/uncertain time information preventing gold standard assignment. Non-CHD primary reason for admission after screening (for MIMIC-IV cohort). \-

Design outcomes

Primary

MeasureTime frameDescription
Overall classification accuracy1 monthTime Frame: Up to completion of dataset evaluation (internal + external cohorts) Description: Proportion of cases with correct subtype (STEMI/NSTEMI/UA/CCS) compared with expert-adjudicated gold standard.

Contacts

CONTACTXiaojin Gao, Dr
sophie_gao@sina.com+86010 88322415

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Mar 5, 2026