Skip to content

Simulated and Synthetic Health Data: Improving Clinical Research on Rare Diseases. A Real-World Data Simulation of Autosomal Dominant Polycystic Kidney Disease (ADPKD) Trials. A Retrospective, Observational Study

Simulated and Synthetic Health Data: Improving Clinical Research on Rare Diseases. A Real-World Data Simulation of Autosomal Dominant Polycystic Kidney Disease (ADPKD) Trials. A Retrospective, Observational Study

Status
Active, not recruiting
Phases
Unknown
Study type
Observational
Source
ClinicalTrials.gov
Registry ID
NCT07016282
Acronym
SAILING-ADPKD
Enrollment
100
Registered
2025-06-11
Start date
2025-06-12
Completion date
2027-06-30
Last updated
2025-09-02

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Autosomal Dominant Polycystic Kidney Disease (ADPKD)

Keywords

rare diseases, RCTs, Monte Carlo simulation, generative AI, ADPKD

Brief summary

This is a no-profit, retrospective observational study involving real-world data (RWD), retrieved from ADPKD-related electronic health records stored at Mario Negri Institute IRCCS. RWD will be used to generate simulated and synthetic datasets, using AI tools. RWD and generated data (GD) will be used to conduct three virtual RCTs, which main outcome is change in Total Kidney Volume (TKV). Statistical tests will be performed to assess quality and privacy preservation of GD compared with RWD. GD will be also evaluated in exploratory sample size estimations.

Detailed description

Randomized clinical trials (RCTs) can be regarded as the least biased source of information to address intervention questions. One of the most common problems encountered in clinical trials focused on rare diseases is the difficulty in finding patients and therefore in building trials on sufficiently large population, in order to have more robust data and less methodological distortions. Several stratagems are already in use to deal with these problems, including extended trial duration, repeated outcome measures, patients genetic profiles, surrogate endpoint, multicenter studies. Another approach is to consider other trial designs in addition to parallel-arms design, such as crossover trial, n-of-1 trials, and adaptive trials. Simulated and synthetic health data can represent new valid approaches to increase the representativeness of the patients, especially in rare diseases field, while reducing costs and time constraints, but also facing the limitations imposed by national and international regulations concerning privacy and data management. Simulation studies are defined as computer experiments that involve creating data by pseudo-random sampling from known probability distributions, based on Monte Carlo method. A promising approach now under development includes synthetic data, defined as artificially generated data with the aim of reproducing the statistical properties of an original dataset, through generative large languages models (LLMs). Thus, while simulated data rely on known distributions that must be specified in advance, synthetic data are generated by LLMs that learn these distributions from training data, without the need for predefined distributions, offering a significant advantage in flexibility and applicability. This study aims to find the most suitable tool for generating simulated and synthetic data in rare diseases field, and to compare the fidelity, quality, and privacy preservation of these datasets, derived from real-world ADPKD clinical trial data. Furthermore, a virtual clinical trial will be conducted using these three datasets to assess their validity in replicating real trial outcomes. Finally, retrieved and generated data will be used to assess new sample size estimations for future clinical trial performed at the Clinical Research Center for Rare Disease Aldo e Cele Daccò, Istituto di Ricerche Farmacologiche Mario Negri IRCCS, Ranica (BG), Italy. By using generative AI models, such as Generative Adversarial Networks (GANs), this study aims to overcome challenges related to data poverty and trial design. The results could provide valuable insights into whether synthetic data can be a useful tool for improving clinical trials in rare diseases, making them more efficient and cost-effective.

Interventions

None listed

Sponsors

Mario Negri Institute for Pharmacological Research
Lead SponsorOTHER

Study design

Observational model
CASE_CONTROL
Time perspective
RETROSPECTIVE

Eligibility

Sex/Gender
ALL
Age
18 Years to No maximum
Healthy volunteers
No

Inclusion criteria

* Adult (\>18 years) men and women with ADPKD according to Ravine criteria25 * Estimated glomerular filtration rate (eGFR) between 15 and 40 mL/min/1.73 m2 (CKD stage: G3b-G4) or higher (CKD stage: G1-G3a), as calculated by the Modification of Diet in Renal Disease study four variables equation

Exclusion criteria

* confounding factors that could affect renal function loss independent of kidney growth and treatment allocation (i.e., diabetes mellitus, urinary protein excretion rate \>3 g/24 h) * Abnormal urinalysis suggestive of concomitant, clinically significant glomerular disease, and urinary tract lithiasis or infection * Patients with major systemic disease * Patients unable to provide informed consent * Pregnant, lactating, or potentially childbearing women without adequate contraception

Design outcomes

Primary

MeasureTime frameDescription
Changes in total kidney volume (TKV)At baseline, and immediately after data generation procedure.Changes in TKV in mL.

Countries

Italy, Sweden

Outcome results

None listed

Source: ClinicalTrials.gov · Data processed: Feb 4, 2026