Medical Image Integrity Verification
Conditions
Keywords
Artificial intelligence, Medical imaging, Synthetic medical image, Patient-information-discordant medical images
Brief summary
This prospective randomized controlled trial will evaluate whether DeepMedFake improves the identification of medical images that require further verification. DeepMedFake is an artificial intelligence-based medical image analysis system that provides an image-authenticity risk assessment, a patient-image consistency risk assessment and a recommendation regarding further verification. Hospital physicians and healthcare audit professionals will participate in a randomized two-period crossover design. Each participant will assess one case set with DeepMedFake assistance and the other without AI assistance. The primary outcome in each cohort is the sensitivity of the final verification decision, evaluated separately in the hospital clinical and healthcare audit cohorts.
Interventions
DeepMedFake is an artificial intelligence-based medical image analysis system designed to support the detection of synthetic medical images and authentic medical images paired with discordant patient information. It provides an image-authenticity risk assessment, a patient-image consistency risk assessment and a recommendation regarding further verification. In this study, DeepMedFake is used as a decision-support tool to assist participants in determining whether the image requires further verification before downstream clinical or healthcare audit use. Final judgments and verification decisions are made by the participants.
Sponsors
Study design
Eligibility
Inclusion criteria
* Aged 18 years or older. * Able to complete reading periods and all required electronic study procedures. * Completed the standardized study training and practice cases. * Provided written informed consent before participation. (1)Hospital clinical cohort participants must also meet the following criteria: * Hold a valid physician qualification. * Currently practice in ophthalmology, radiology, ultrasound medicine, or a clinical specialty in which imaging of the neurological or cerebrovascular, cardiovascular, thoracic, abdominal, or musculoskeletal domain is routinely reviewed. * Have direct professional experience reviewing all imaging modalities and clinical imaging domains assigned to them in the study. * Routinely use the assigned medical images for image interpretation, image verification or clinical decision-making. (2)Healthcare audit cohort participants must also meet the following criteria: * Currently work in healthcare audit, medical reimbursement review, medical-cost review or medical-material verification. * Have experience reviewing medical imaging materials as part of their routine work. * Have knowledge required to interpret the medical images and patient information presented in the study and to determine whether further verification is required.
Exclusion criteria
* Direct involvement in development of the locked DeepMedFake model, determination of model weights or selection of decision thresholds. * Direct involvement in selection or construction of the formal case library or adjudication of the reference standard. * Direct involvement in generation or implementation of the reader randomization sequence. * Previous participation as a reader in the pilot study. * Previous access to any formal study case, case-construction record or reference-standard label. * Access to undisclosed study information that could permit advance identification of case type or image-generation method. * A financial, professional or other conflict of interest considered likely to compromise independent case assessment.
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Reader Performance With and Without DeepMedFake Assistance: Sensitivity of the Final Verification Decision in the Hospital Clinical Cohort | During both reading periods, up to 4 weeks after randomization | The percentage of reference-standard abnormal cases assigned to further verification by hospital physicians will be estimated under DeepMedFake-assisted and unassisted review. The effect measure is the absolute percentage-point difference between review conditions. |
| Reader Performance With and Without DeepMedFake Assistance: Sensitivity of the Final Verification Decision in the Healthcare Audit Cohort | During both reading periods, up to 4 weeks after randomization | The percentage of reference-standard abnormal cases assigned to further verification by healthcare audit professionals will be estimated under DeepMedFake-assisted and unassisted review. The effect measure is the absolute percentage-point difference between review conditions. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Reader Performance With and Without DeepMedFake Assistance: Specificity of the Final Verification Decision in the Hospital Clinical Cohort | During both reading periods, up to 4 weeks after randomization | The percentage of authentic, information-matched cases assigned to no further verification by hospital physicians will be estimated under each review condition and compared between conditions. |
| Reader Performance With and Without DeepMedFake Assistance: Specificity of the Final Verification Decision in the Healthcare Audit Cohort | During both reading periods, up to 4 weeks after randomization | The percentage of authentic, information-matched cases assigned to no further verification by healthcare audit professionals will be estimated under each review condition and compared between conditions. |
| Sensitivity of the Final Verification Decision for Synthetic Medical Image Cases | During both reading periods, up to 4 weeks after randomization | The percentage of reference-standard synthetic-image cases assigned to further verification will be estimated under each review condition. DeepMedFake-assisted and unassisted review will be compared separately within the hospital clinical and healthcare audit cohorts. |
| Sensitivity of the Final Verification Decision in Discordant Cases | During both reading periods, up to 4 weeks after randomization | The percentage of authentic images paired with discordant patient information that are assigned to further verification will be estimated under each review condition. DeepMedFake-assisted and unassisted review will be compared separately within the hospital clinical and healthcare audit cohorts. |
| Sensitivity of the Image-Authenticity Judgment | During both reading periods, up to 4 weeks after randomization | The percentage of reference-standard synthetic-image cases classified as suspected synthetic images will be estimated under each review condition and compared separately within each reader cohort. |
| Specificity of the Image-Authenticity Judgment | During both reading periods, up to 4 weeks after randomization | The percentage of reference-standard authentic-image cases classified as not suspected to be synthetic will be estimated under each review condition and compared separately within each reader cohort. |
| Sensitivity of the Patient-Image Consistency Judgment | During both reading periods, up to 4 weeks after randomization | Among authentic images, the percentage of discordant cases classified as discordant with the displayed patient information will be estimated under each review condition and compared separately within each reader cohort. Synthetic-image cases will be excluded from this analysis. |
| Specificity of the Patient-Image Consistency Judgment | During both reading periods, up to 4 weeks after randomization | Among authentic images, the percentage of information-matched cases classified as matching the displayed patient information will be estimated under each review condition and compared separately within each reader cohort. Synthetic-image cases will be excluded from this analysis. |
| Case-Level Workflow Time | During both reading periods, up to 4 weeks after randomization | Elapsed time in seconds from successful case loading to submission of the final required case-level response, including selection of a proposed first verification pathway when required. Recorded pauses will be excluded. Workflow time will be compared between review conditions separately within each reader cohort. |
Countries
China