Skip to content

Prospective Evaluation of the Accuracy of Four Large Language Model Chatbots in Answering Pediatric Surgical Disease-Related Questions: An Experimental Comparative Study

Prospective Evaluation of the Accuracy of Four Large Language Model Chatbots in Answering Pediatric Surgical Disease-Related Questions: An Experimental Comparative Study

Status
Active, not recruiting
Phases
Unknown
Study type
Observational
Source
ChiCTR
Registry ID
ChiCTR2600117120
Enrollment
Unknown
Registered
2026-01-20
Start date
2026-01-23
Completion date
Unknown
Last updated
2026-01-27

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Pediatric Urinary System Disorders

Interventions

Chatgpt-4o group:None
Gemini2.5pro group:None
Openevidence group:None
Zhipu Qingyan (ChatGLM) group:None

Sponsors

West China Hospital of Sichuan University
Lead Sponsor

Eligibility

Sex/Gender
All

Inclusion criteria

Inclusion criteria: 1. Disease and Patient Inclusion Criteria: (1) Clearly diagnosed with a specific common pediatric surgical disease (such as pediatric urological diseases or general pediatric surgical diseases); (2) The patient’s age falls within the high-risk group for the disease, with stable condition and no severe complications; (3) The patient’s family is willing to participate in an open-ended questionnaire survey and can clearly express concerns related to the disease. 2. Research Participant Inclusion Criteria: (1) At least 18 years old, possessing full civil capacity, and a direct relative of the selected patient (parents, grandparents); (2) Have basic reading comprehension and expression abilities, able to independently complete the questionnaire (provide questions, rank order); (3) Understand the purpose of this study, voluntarily participate, and sign a written informed consent form; (4) No communication barriers and able to cooperate with the research team in collecting and ranking questions. 3. AI Chatbot Inclusion Criteria: (1) Mainstream AI chatbots, including ChatGPT 4 (OpenAI), Copilot, Claude AI, and Baidu's Wenxin Yiyan; (2) Includes both free and paid versions, and all must be the latest version at the start of the study; (3) Must support text input and response output, and can consistently provide response results. 4. Scoring Expert Inclusion Criteria: (1) Belongs to a field related to the diseases involved in the study; (2) Familiar with diagnostic and treatment guidelines and authoritative medical literature on the disease; (3) No conflicts of interest and voluntarily participate in the scoring work.

Exclusion criteria

Exclusion criteria: 1. Disease and Pediatric Patient Exclusion Criteria: (1) Children with rare pediatric surgical diseases, multiple system malformations, or critical conditions (such as severe infection, shock, etc.); (2) Children with other serious organic diseases or psychiatric disorders that may affect family members' understanding of the disease; (3) Children with unclear disease diagnosis or undetermined treatment plans. 2. Research Participant Exclusion Criteria: (1) Under 18 years old or without full civil capacity; (2) Having cognitive impairment, psychiatric illness, or communication barriers that prevent clear expression of problems or completion of questionnaires; (3) Refusing to sign the informed consent form or not cooperating in the collection and prioritization of questions; (4) Having a conflict of interest with the research team members, which may affect the objectivity of the questions provided; (5) Participated in similar AI chatbot evaluation studies within the past 3 months. 3. AI Chatbot Exclusion Criteria: (1) AI chatbots that cannot provide stable response results; (2) AI chatbots without a clear distinction between free and paid versions; (3) Outdated AI chatbot versions that have been discontinued or stopped updating at the start of the study; (4) AI chatbots with technical barriers that prevent inputting questions in a standardized format. 4. Scoring Expert Exclusion Criteria: (1) Pediatric surgeons not related to the disease area; (2) Having a conflict of interest (e.g., a cooperative relationship with AI chatbot development companies); (3) Refusing to comply with blinded scoring principles or not cooperating in the scoring work.

Design outcomes

Primary

MeasureTime frame
AI Chatbot Response Quality;

Secondary

MeasureTime frame
Questionnaire Reliability and AI Response Acceptance;

Countries

China

Contacts

Public ContactKang Ting

West China Hospital of Sichuan University

Kt0407hyll@163.com+86 177 3683 5867

Outcome results

None listed

Source: ChiCTR (via WHO ICTRP) · Data processed: Feb 4, 2026