Skip to content

AI vs Human Exam Assessment and Development (AHEAD Trial)

Psychometric Performance and Student Perceptions of AI- Versus Human-Generated Multiple-Choice Question Development in Medical Education: The AHEAD Randomized Controlled Trial

Status
Completed
Phases
Unknown
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT07481162
Acronym
AHEAD
Enrollment
258
Registered
2026-03-18
Start date
2024-12-08
Completion date
2024-12-09
Last updated
2026-03-18

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Medical Education Assessment

Keywords

Artificial Intelligence, Medical Education, Large Language Models, Multiple Choice Questions, Educational Assessment, Randomized Controlled Trial, ChatGPT

Brief summary

The Artificial Intelligence (AI) vs Human Exam Assessment and Development (AHEAD) Trial is a participant-blinded randomized controlled trial conducted among first-year medical students at the University of British Columbia. The study evaluates whether multiple-choice examination questions generated using large language models (LLMs) perform comparably to traditionally human-written questions in medical education. Participants were randomized to complete one of two versions of a formative mock final examination consisting of 112 case-based single-best-answer multiple-choice questions (MCQs) aligned with the same course learning objectives. One exam version contained AI-generated questions produced using a structured LLM workflow with independent AI verification, while the other contained questions authored by senior medical students using conventional methods. The study evaluates exam feasibility, psychometric reliability, validity, student acceptability, and educational impact. Outcomes include exam performance, item discrimination indices, distractor efficiency, student perceptions of exam quality and difficulty, and changes in perceived preparedness for the upcoming summative examination.

Detailed description

The AHEAD Trial (AI vs Human Exam Assessment and Development) is a single-center, participant-blinded randomized controlled trial conducted among first-year Doctor of Medicine (MD) students enrolled in the Foundations of Medical Practice I (MEDD 411) course at the University of British Columbia. Participants were randomized in a 1:1 ratio to complete either an AI-generated or a human-generated mock final examination. Both exams consisted of 112 case-based single-best-answer multiple-choice questions (MCQs) aligned with the same MEDD 411 curricular objectives. AI-generated questions were produced using a structured workflow involving ChatGPT for question generation and Google Gemini for independent verification. Human-generated questions were authored by senior medical students without AI assistance and underwent independent peer review. Both exams followed identical formatting guidelines and assessed the same learning objectives. All participants completed identical pre-exam and post-exam surveys assessing demographic characteristics, familiarity with artificial intelligence in education, and perceptions of the examination experience. The study evaluates the utility of AI-generated assessments using van der Vleuten's Assessment Utility Framework, including feasibility, reliability, validity, acceptability, and educational impact. The trial aims to determine whether large language models can accelerate the development of formative medical examinations while maintaining comparable psychometric quality and educational value relative to traditional human-authored questions.

Interventions

OTHERAI-generated MCQ examination

A formative mock examination composed of 112 case-based multiple-choice questions generated using large language models aligned with course learning objectives.

OTHERHuman-generated MCQ examination

A formative mock examination composed of 112 case-based multiple-choice questions written by senior medical students using conventional item-writing methods aligned with the same course learning objectives.

Sponsors

University of British Columbia
Lead SponsorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
PARALLEL
Primary purpose
OTHER
Masking
SINGLE (Subject)

Masking description

Participants were blinded to the source of the examination questions (AI-generated vs human-generated). All items were reviewed to remove indicators of authorship before distribution.

Intervention model description

Participants were randomized 1:1 to complete either an AI-generated or human-generated mock examination composed of 112 case-based multiple-choice questions aligned with the same course learning objectives.

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

* Enrolled first-year medical students in the University of British Columbia MD undergraduate program. * Can voluntarily consent to participate in the formative mock examination study.

Exclusion criteria

* Students who declined participation. * Students who did not complete the mock examination or required surveys.

Design outcomes

Primary

MeasureTime frameDescription
Student performance on the mock examinationImmediately after completion of the mock examinationComparison of mean examination scores between students randomized to the AI-generated versus human-generated mock examinations.

Secondary

MeasureTime frameDescription
Item discrimination indexImmediately after the completion of the mock examinationItem-level discrimination index comparing AI-generated and human-generated multiple-choice questions, representing the difference in the proportion of correct responses between high-performing and low-performing students.
Distractor efficiencyImmediately after the completion of the mock examinationProportion of distractors selected by at least 5% of participants, comparing AI-generated and human-generated questions.
Student-rated examination quality and acceptabilityImmediately after completion of the mock examinationStudent ratings of exam difficulty, clarity, relevance to course material, adequacy of time, multiple-choice question quality, understanding of clinical concepts, identification of knowledge gaps, retention for future clinical practice, and preparedness for the upcoming summative exam, measured immediately after exam completion using 10-point Likert scales (1 = lowest rating, 10 = highest rating). For most domains, higher scores indicate greater endorsement of the construct being measured; for the difficulty item, higher scores indicate greater perceived difficulty.
Efficiency ratio of MCQ development time per matched learning objectiveBaseline (prior to participant testing)The outcome measuring the development efficiency of artificial intelligence (AI)-generated versus human-generated multiple-choice questions (MCQs). The efficiency ratio was calculated as human-generated MCQ development time divided by AI-generated MCQ development time for matched learning objectives.
Change in perceived preparedness for the summative examinationBefore and immediately after completion of the mock examinationChange from pre-exam to post-exam in self-rated preparedness for the upcoming summative examination, measured on a 10-point Likert scale (1 = not at all prepared; 10 = extremely prepared), with higher scores indicating greater perceived preparedness.

Countries

Canada

Contacts

PRINCIPAL_INVESTIGATORAnita Palepu, MD, MPH, FRCPC

University of British Columbia

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Jun 19, 2026