Skip to content

Which works better - summaries of health research written by people alone or with computer help?

Comparison of AI-assisted and human-generated plain language summaries for Cochrane reviews: protocol for a randomised trial

Status
Active, not recruiting
Phases
Unknown
Study type
Interventional
Source
ISRCTN
Registry ID
ISRCTN85699985
Enrollment
454
Registered
2025-02-04
Start date
2025-09-01
Completion date
Unknown
Last updated
2026-02-02

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Health information Other

Interventions

This study is a randomised, parallel-group, two-armed, non-inferiority trial designed to compare the effectiveness of AI-assisted Cochrane plain language summarises (PLS) and standard human-generated

Sponsors

Ollscoil na Gaillimhe – University of Galway
Lead Sponsor

Eligibility

Sex/Gender
All
Age
18 Years to 100 Years

Inclusion criteria

Inclusion criteria: 1. Participants must be 18 years or older. 2. Participants must be proficient in English, as the study materials-including the intervention, comparator, and assessments—will be provided in English. To ensure adequate comprehension, participants will be asked to self-report their English reading proficiency on a scale from 1 (very poor) to 10 (excellent), and only those who rate their reading proficiency as 7 or higher will be eligible to participate in the study. 3. Participants must have access to the internet and a device capable of completing an online survey (e.g., computer, tablet, or smartphone). 4. Participants must provide informed consent before starting the study.

Exclusion criteria

Exclusion criteria: 1. Individuals with formal education or professional experience in health-related fields, such as healthcare professionals or academic researchers in health or medicine, will be excluded. This ensures the study targets individuals who may have less familiarity with the topic, thereby maximising the potential impact of the findings. 2. Individuals unable to complete the online survey or who fail to meet the minimum participation requirements will be excluded. 3. Responses will be excluded if they meet any of the following criteria: 3.1. Total completion time <10 minutes per summary (combined reading and question time) 3.2. Evidence of straight-line answering patterns 3.3. Inconsistent answers to related questions

Design outcomes

Primary

MeasureTime frame
1. Comprehension measured using a standardised 10-item multiple-choice questionnaire for each summary at one timepoint 2. Readability measured using the Flesch-Kincaid Grade Level at one timepoint

Secondary

MeasureTime frame
1. Readability measured using the following at one timepoint: 1.1. Flesch-Kincaid Grade Level 1.2. Gunning Fog Index 1.3. Automated Readability Index (ARI) 1.4. Coleman-Liau Index 1.5. SMOG Index 1.6. Linsear Write Readability Formula 2. Quality of information will be assessed by two independent systematic review experts using a standardised assessment form and rating scale at one timepoint focusing on four main types of errors: 2.1. Incorrect output (where the LLM generated wrong information) 2.2. Irrelevant output (where the LLM generated unnecessary or off-topic information) 2.3. Omissions (where the LLM failed to include key information that should be present) 2.4. Currency errors (where information is outdated or inconsistent with current evidence) 3. Safety measured by expert raters who will evaluate each summary for potential risks and biases at one timepoint, focusing on: 3.1. Risk of misinterpretation 3.2. Presence of bias or inappropriate recommendations 3.3. Appropriate presentation of limitations and uncertainties 3.4. Evidence of fabrication or hallucination 4. Perceived trustworthiness by asking participants to assess the trustworthiness of the reviews using a 5-point Likert scale based on items adapted from existing scales measuring trust in online health information. Participants will rate their agreement with the following statements: 4.1. "I trust the information provided in this summary." 4.2. "This summary is from a reliable source." 4.3. "I am confident in the accuracy of the information in this summary." 4.4. "I believe the source of this summary has expertise in the subject matter." 4.5. "I would use the information from this summary to make health decisions." After completing all assessments for all summaries, participants will be asked two additional questions: 1. Do you think this summary was written by: 1.1. A human expert 1.2. An AI system with human expert review 1.3. Not sure 2. How much would it matter to you if health info

Countries

United States of America

Contacts

Public ContactDeclan Devane
declan.devane@universityofgalway.ie+353 91 524411

Outcome results

None listed

Source: ISRCTN (via WHO ICTRP) · Data processed: Feb 7, 2026