Skip to content

Specialist Review of Harm Ratings for Clinical Responses Acronym: NOHARM

Specialist Physician Validation of the NOHARM Harm-Rating Rubric for Clinical Responses

Status
Not yet recruiting
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07823062
Enrollment
80
Registered
2026-09-16
Start date
2026-09-21
Completion date
2026-10-30
Last updated
2026-09-16

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Medical Errors, Patient Safety

Keywords

Artificial Intelligence, Clinical Response Evaluation, Harm Assessment, Clinical Benchmark, Rubric Validation

Brief summary

The purpose of this study is to determine how closely potential-harm ratings assigned using the NOHARM rubric agree with independent ratings from specialty-matched physicians. The study uses reconstructed clinical consultation cases and previously generated responses written by artificial intelligence systems or physicians. The materials contain no patient identifiers or protected health information. Board-certified attending physicians from 10 medical specialties will review assigned case-response pairs through a secure website. For each response, they will indicate whether it contains an error, describe the error when applicable, and rate the greatest potential patient harm as none, mild, moderate, or severe. Reviewers will not be shown the response source, existing rubric rating, or other reviewers' ratings. Two specialists will initially review each response; a third may review it if the first two ratings differ. The study includes a development stage, possible rubric revision, and a separate validation stage.

Interventions

None listed

Sponsors

Stanford University
Lead SponsorOTHER

Study design

Observational model
COHORT
Time perspective
PROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
Yes

Inclusion criteria

* Licensed or registered physician in good standing * Attending- or consultant-level physician actively practicing in allergy and immunology, cardiology, dermatology, endocrinology, gastroenterology, hematology, infectious disease, nephrology, neurology, or pulmonology * Practicing in the United States, Canada, United Kingdom, or Australia * Assigned clinical materials match the physician's specialty * Able to read and respond in English * Access to a computer with a stable internet connection * Willing and able to complete at least one remote specialist-rating engagement of approximately three hours

Exclusion criteria

* Does not meet the professional, career-level, specialty, geographic, or credential-verification requirements above * Prior involvement with assigned cases, responses, or rubric information that could compromise independent rating * Unable to complete the remote rating activities in English * Unable or unwilling to follow the independent-rating procedures * For validation-stage assignments, participation in the rubric-revision process

Design outcomes

Primary

MeasureTime frameDescription
Inverse-Probability-Weighted Confirmation of Rubric-Severe ClassificationsAt completion of the development and validation stages, up to 2 monthsThe inverse-probability-weighted proportion of responses classified as severe by the NOHARM rubric that are also classified as severe by specialist consensus: P(specialist severe \| rubric severe). This measure will be calculated separately for the development and validation stages. Values range from 0 to 1; higher values indicate greater confirmation. The prespecified threshold is 0.70.
Inverse-Probability-Weighted Capture of Specialist-Severe ClassificationsAt completion of the development and validation stages, up to 2 monthsThe inverse-probability-weighted proportion of responses classified as severe by specialist consensus that are also classified as severe by the NOHARM rubric: P(rubric severe \| specialist severe). This measure will be calculated separately for the development and validation stages. Values range from 0 to 1; higher values indicate greater capture. The prespecified threshold is 0.70. Both co-primary measures must meet this threshold.

Contacts

CONTACTJonathan H Chen, MD, PHD
jonc101@stanford.edu650-725-3655

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Sep 17, 2026