Skip to content

Assessing the Performance of Artificial Intelligence (AI)-Augmented Electronic Health Record (EHR) Data Abstraction for Clinical Trial Patient Screening

Assessing the Performance of Artificial Intelligence (AI)-Augmented Electronic Health Record (EHR) Data Abstraction for Clinical Trial Patient Screening

Status
Completed
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT06561217
Enrollment
355
Registered
2024-08-19
Start date
2023-08-18
Completion date
2024-07-12
Last updated
2025-07-25

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Cancer

Brief summary

Identifying eligible patients is a key process in the clinical trial enterprise. Currently, this process relies on time-intensive manual chart review, creating a rate-limiting step for trial participation. The integration of AI technology into the trial screening process has potential to improve participation rates. This study aims to assess the performance (accuracy, efficiency) of AI-augmented patient identification and inform optimal integration into clinical research screening processes.

Detailed description

The objective of this study is to assess and compare the accuracy and efficiency of three different approaches to abstracting clinical data used to identify oncology patients who meet the inclusion criteria for participation in clinical trials. The three approaches under evaluation include: (1) an autonomous AI algorithm (Mendel AI; developed by artificial intelligence startup company Mendel) which analyzes patient medical records to extract relevant clinical facts (AI-alone); (2) a human researcher who manually reviews patient charts as per the current norm/practice (Human-alone); and (3) a human researcher utilizing AI augmentation (Human+AI), where Mendel AI serves as a supportive tool in the decision-making process by providing the researcher a list of elements abstracted by the AI algorithm and a rank-order list of patients most likely to meet inclusion criteria for a trial. The study primarily aims to compare (1) the chart-level accuracy of the Human+AI collaboration relative to Human-alone given the relevance of this comparison for real-world clinical workflows, defined by the percentage of pre-identified chart elements classified correctly compared against a predetermined gold standard; and (2) the efficiency of the Human+AI vs. Human-alone arms, defined by the time per chart review in minutes, measured for each chart. Our hypotheses are (1) the Human+AI arm will be non-inferior in accuracy when compared to the Human-alone arm, in relation to a predetermined gold standard, and (2) that a Human+AI arm will be superior in efficiency of abstraction when compared to Human-alone screening. The identification of eligible patients for clinical trials is a critical component of clinical research, as it directly impacts patient recruitment, study enrollment, and the generalizability of research findings. Currently, the process of identifying eligible patients often relies on manual chart review by clinical research staff, which can be time-consuming, labor-intensive, and prone to human error. Consequently, eligible patients may be overlooked, and opportunities for trial participation may be missed. The integration of AI technology into the patient identification process has the potential to enhance the accuracy and efficiency of this critical task, leading to improved clinical trial recruitment and outcomes. This study holds important implications for the field of clinical research by evaluating the effectiveness of AI-augmented patient identification compared to traditional manual methods and autonomous AI algorithms. By examining the strengths and limitations of each approach, the study will provide valuable insights into the optimal integration of AI technology in clinical research processes. Furthermore, the results of this study have the potential to benefit patients by improving their access to clinical trials and increasing awareness of available treatment options. For clinical research institutions, enhancing the efficiency of patient identification can lead to more effective use of research resources and the potential for accelerated clinical trial timelines. Ultimately, the findings of this study may contribute to advancements in clinical research practices, promoting more equitable access to trials and facilitating the development of innovative treatments for patients with cancer.

Interventions

OTHERChart review

Chart review

Sponsors

Mendel AI
CollaboratorINDUSTRY
University of Pennsylvania
Lead SponsorOTHER

Study design

Observational model
OTHER
Time perspective
RETROSPECTIVE

Eligibility

Sex/Gender
ALL
Healthy volunteers
No

Inclusion criteria

* Diagnosis of colorectal or non-small cell lung cancer. * A minimum of 5 patient documents in the Mendel database. * Most recent document was within 5 years from the time of data extraction.

Exclusion criteria

* None.

Design outcomes

Primary

MeasureTime frameDescription
Abstracted Chart-level Accuracy1 yearThe primary outcome measured was mean chart-level accuracy, defined as the percentage of elements identified by clinical research coordinators among all elements in the gold-standard set, measured for each chart, and averaged across all charts. Research coordinator-abstracted responses were identified as being accurate when they exactly matched with the gold-standard set. The gold-standard set was determined by 2-3 clinicians blinded to experimental arms.

Secondary

MeasureTime frameDescription
Efficiency of Chart-level Abstraction (in Minutes)1 yearEfficiency was calculated as the number of minutes spent on each chart abstraction.

Countries

United States

Participant flow

Pre-assignment details

Each record will be reviewed via all 3 arms, so the total enrollment will be 355.

Participants by arm

ArmCount
All Participants
Patients who underwent Human-alone, AI-alone, and Human+AI chart review
355
Total355

Baseline characteristics

CharacteristicAll Participants
Age, Customized
<18 years
0 Participants
Age, Customized
≥18 years
355 Participants
Cancer Type
Colorectal Cancer
160 Participants
Cancer Type
Non-Small Cell Lung Cancer
195 Participants
Race and Ethnicity Not Collected— Participants
Region of Enrollment
United States
355 participants
Sex/Gender, Customized
Any sex/gender
— Participants

Adverse events

Event typeEG000
affected / at risk
EG001
affected / at risk
EG002
affected / at risk
deaths
Total, all-cause mortality
0 / 00 / 00 / 0
other
Total, other adverse events
0 / 00 / 00 / 0
serious
Total, serious adverse events
0 / 00 / 00 / 0

Outcome results

Primary

Abstracted Chart-level Accuracy

The primary outcome measured was mean chart-level accuracy, defined as the percentage of elements identified by clinical research coordinators among all elements in the gold-standard set, measured for each chart, and averaged across all charts. Research coordinator-abstracted responses were identified as being accurate when they exactly matched with the gold-standard set. The gold-standard set was determined by 2-3 clinicians blinded to experimental arms.

Time frame: 1 year

Population: The study population was drawn from a 15-physician community oncology practice in California serving patients from urban and surrounding rural communities. This study cohort consisted of unstructured medical records from patients within the dataset with 1) a diagnosis of non-small cell lung cancer (NSCLC) or colorectal cancer (CrCa), 2) a minimum of five clinical documents available, and 3) the most recent document being within five years from the time of data extraction.

ArmMeasureValue (MEAN)Dispersion
Human + AIAbstracted Chart-level Accuracy76.10 % of elements correctly abstractedStandard Deviation 20.59
Human-aloneAbstracted Chart-level Accuracy71.48 % of elements correctly abstractedStandard Deviation 24.92
AI-aloneAbstracted Chart-level Accuracy59.92 % of elements correctly abstractedStandard Deviation 23.75
Comparison: A Shapiro-Wilk test was used to assess normality of paired differences in chart-level accuracy (alpha = 0.05). A one-sided, paired Wilcoxon Rank Sum test (alpha = 0.05) was employed to test the primary null hypothesis of whether chart-level accuracy of the Human+AI arm for EHR chart abstraction was non-inferior to the chart-level accuracy of a Human-alone arm abstraction by at least 5% (i.e., noninferiority margin).p-value: <0.001Wilcoxon (Mann-Whitney)
Comparison: A Shapiro-Wilk test was used to assess normality of paired differences in chart-level accuracy (alpha = 0.05). A one-sided, paired Wilcoxon Rank Sum test (alpha = 0.05) was employed to test the null hypothesis of whether chart-level accuracy of the Human+AI arm for EHR chart abstraction was superior to the chart-level accuracy of a Human-alone arm abstraction.p-value: 0.002Wilcoxon (Mann-Whitney)
Secondary

Efficiency of Chart-level Abstraction (in Minutes)

Efficiency was calculated as the number of minutes spent on each chart abstraction.

Time frame: 1 year

Population: The study population was the same as the study population described in the primary outcome analysis population description section. However, AI-alone chart reviews were NOT analyzed for efficiency. Data for the secondary outcome of efficiency was not collected for the AI-alone arm and therefore cannot be reported in the outcome table. This was prespecified, as its fully automated nature renders direct comparison with human-involved workflows inappropriate and uninformative.

ArmMeasureValue (MEDIAN)
Human + AIEfficiency of Chart-level Abstraction (in Minutes)32.12 Number of minutes spent on abstraction
Human-aloneEfficiency of Chart-level Abstraction (in Minutes)31.75 Number of minutes spent on abstraction
Comparison: A two-sided, paired Wilcoxon Rank Sum test (alpha = 0.05) was employed to test for difference between chart-level efficiency of the Human+AI arm and chart-level efficiency of the Human-alone arm abstraction.p-value: 0.51395% CI: [-1.25, 2.55]Wilcoxon (Mann-Whitney)

Source: ClinicalTrials.gov · Data processed: Feb 4, 2026