Bladder Cancer, Prostate Cancer, Urinary Stones
Conditions
Keywords
Urological Diseases, Urinary Stones, Prostate Cancer, Bladder Cancer, Artificial Intelligence, Large Language Model, AI Agent, Clinical Decision Support
Brief summary
Urological diseases such as urinary stones, prostate cancer, and bladder cancer are very common and often require highly specialized diagnosis and treatment. Today, the quality of care can vary between doctors, and there are not enough urology specialists to meet patient demand. Artificial intelligence (AI) may help doctors make faster and more consistent decisions. This study aims to develop and test an AI-powered assistant called "UroAgent" that supports doctors in diagnosing and treating urological diseases. UroAgent is built on a large language model trained specifically for urology and is connected to tools that help it retrieve medical knowledge and analyze images. To build and test UroAgent, the research team will use 1,500 past patient records from 2010-2025 and collect 500 new patient cases for validation, for a total of 2,000 cases. This is an observational study: no patient's medical treatment will be changed because of it. The goal is to create a reliable AI tool that helps improve urological care for patients.
Detailed description
This study protocol describes an observational study aiming to develop and validate UroAgent, an artificial-intelligence agent for the diagnosis and treatment of urological diseases. A total of 2,000 urological disease cases will be collected, comprising 1,500 retrospective cases recorded at the center between 2010 and 2025 for model development and 500 prospectively enrolled cases for independent performance validation. The primary evaluation is the concordance between UroAgent's diagnostic and treatment recommendations and the reference standards established by senior urologists, assessed through diagnostic accuracy, recommendation appropriateness, completeness, and safety; secondary evaluations include the agent's performance across disease subtypes (urinary stones, prostate cancer, bladder cancer) and its image-interpretation capability. All records will undergo de-identification, and the study will adhere to rigorous ethical standards and a pre-specified statistical analysis plan to provide robust evidence for the clinical application of this urology-specific AI agent.
Interventions
None listed
Sponsors
Study design
Eligibility
Inclusion criteria
1. Diagnosed with a urological disease (e.g., urinary stones, prostate cancer, bladder cancer, and other urological conditions). 2. Availability of complete clinical information, imaging data, and surgical video required for model development and validation.
Exclusion criteria
1\. Missing clinical information, imaging data, or surgical video.
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Diagnostic Accuracy of UroAgent | Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up. | The primary outcome is UroAgent's diagnostic accuracy, measured as the F1 score of its leading diagnosis against the reference-standard final diagnosis. The reference standard is established by senior urologists from pathology, imaging, and clinical course. F1 = 2 × Precision × Recall / (Precision + Recall), computed per case and aggregated as macro-F1 across the 2,000-case cohort (1,500 retrospective + 500 prospective). Unit of measure: F1 score (range 0-1). |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Expert Subjective Accuracy Rating | Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up. | Independent urologists rate the accuracy of UroAgent's diagnosis on a 5-point Likert scale (1 = completely inaccurate, 5 = completely accurate), blinded to model identity. Reported as the mean rating and the percentage of cases rated ≥ 4; disagreements resolved by a third urologist. Unit of measure: mean rating (1-5) and percentage of cases rated accurate (%). |
| Treatment Recommendation Appropriateness of UroAgent | Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up. | The proportion of cases in which UroAgent's management recommendation is rated appropriate. Two independent urologists rate each recommendation on a 5-point Likert appropriateness scale (1 = completely inappropriate, 5 = completely appropriate), blinded; "appropriate" defined as Likert ≥ 4; disagreement resolved by a third urologist. Unit of measure: percentage of cases (%) rated appropriate. |
| Clinical Safety of UroAgent Recommendations | Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up. | The rate of clinically unsafe recommendations. Each case is reviewed by senior urologists using a safety rubric and flagged (binary per case) for any contraindicated, erroneous, or potentially harmful recommendation. Unit of measure: percentage of cases (%) with ≥ 1 unsafe recommendation. |
| Concordance and Non-Inferiority of UroAgent versus Clinician Diagnoses | Retrospective cases - at data extraction (single time point); Prospective cases - at enrollment (single time point); no longitudinal follow-up. | On the same cases, UroAgent's diagnoses are compared with those of practicing urologists. Reported as the agreement rate (%) between UroAgent and clinician diagnoses, plus the diagnostic-accuracy difference in percentage points (pp); Cohen's κ reported as a supplementary statistic. Both are assessed against the reference-standard final diagnosis. Unit of measure: agreement rate (%) and accuracy difference (pp). |
Countries
China