Simulate a trial and compare synthesizers.
Four questions to a starting method.
Fifteen questions with hints.
Slides, tools and glossary.
Simple, transparent and often enough
These methods should be tried first. For trial-sized datasets they are frequently as good as anything more complex.
Family 01 · Rules and simulation
Knowledge-driven simulation
You write down what you believe about the population and the trial, then let the computer generate patients. Statisticians have done this for decades to check sample sizes and trial designs. Examples include a two-arm trial with assumed event rates and dropout, PK/PD models linking dose to concentration and response, and Synthea, which simulates patient records along care pathways.
A two-arm survival trial in about ten lines of Python. Every number is an assumption you can change. Try it interactively in the lab →
import numpy as np, pandas as pd
rng = np.random.default_rng(2026)
n = 300 # patients per arm
median_ctrl, hr = 6.0, 0.75 # assumptions (months, hazard ratio)
lam_c = np.log(2) / median_ctrl # exponential hazard, control
arm = np.repeat(["Placebo", "Active"], n)
lam = np.where(arm == "Active", lam_c * hr, lam_c)
t_evt = rng.exponential(1 / lam) # time to event
t_cen = rng.uniform(6, 24, 2 * n) # administrative censoring
df = pd.DataFrame({"arm": arm,
"time": np.minimum(t_evt, t_cen),
"event": (t_evt <= t_cen).astype(int)})Family 02 · Resampling
Resampling and perturbation
Bootstrap draws real records with replacement. Noise addition makes small random changes so no record is an exact copy. SMOTE-type methods create new records between existing ones, often to boost rare classes.
Family 03 · Statistical models
The Gaussian copula in three steps
- Learn each column. What does age look like on its own? Tumour size? Learn the shape of every variable.
- Learn how they move together. Older patients have more comorbidities. Learn these correlations separately.
- Generate new rows. Draw correlated random numbers and map them back onto each column's shape.
Family 03 · Statistical models
Build a patient one variable at a time
Bayesian networks and sequential CART (for example the R package synthpop) generate each variable from the ones already generated, much like following a clinical story.
It handles mixed data types and nonlinear relationships well, lets you build clinical logic into the variable order, and works with a few hundred to a few thousand records. The order matters, and very many variables make it slow.
library(synthpop)
syn_obj <- syn(real_df, method = "cart", seed = 2026)
synthetic_df <- syn_obj$syn
compare(syn_obj, real_df) # side-by-side distributionsWhen relationships get complex
Family 04 · Deep learning
Why deep learning?
Deep generators help when there are complex, interacting relationships, hundreds of variables, sequences over time such as visits and treatment histories, or unstructured text and images.
Family 04 · Deep learning
GANs: a forger and a detective
A generator (the forger) turns random noise into fake patients. A discriminator (the detective) tries to tell fake from real. Each round of feedback makes both better, and training stops when the detective can no longer tell the difference.
Tabular versions for mixed clinical data include CTGAN, CopulaGAN and Wasserstein GANs such as WGAN-GP and conditional tabular WGANs.
Family 04 · Deep learning
Three more families worth knowing
VAEs (e.g. TVAE)
Compress each patient into a small summary, then learn to rebuild patients from it. Often more stable than GANs.
Diffusion (e.g. TabDDPM)
Add noise step by step, then learn to remove it. Start from noise to get new patients. High quality, slower.
Transformers and LLMs
Treat a patient history as a sentence of events and generate the next event. Promising for EHR sequences; privacy needs care.
Clinical reality
What makes clinical data hard to synthesise
| Challenge | What to check |
|---|---|
| Survival and censoring | Time and event status stay consistent; Kaplan–Meier curves match. |
| Small samples | Trials often have hundreds, not millions, of patients. Prefer stable methods. |
| Rare events | Serious adverse events and small subgroups are still present. |
| Clinical rules | Death after randomisation, doses within range, valid lab units. |
| Mixed data types | Continuous, categorical, dates and text in one dataset. |
| Longitudinal structure | Repeated visits per patient remain coherent over time. |
Case study
Synthetic data from a lung cancer trial
Using a randomised lung cancer trial (Tarceva versus placebo) with overall survival and progression-free survival endpoints, provided by RSS under a data-sharing agreement, we compared a Gaussian Copula Synthesizer with a conditional tabular Wasserstein GAN.
Both were judged with the same checks: Kaplan–Meier curves, hazard ratios, distributions and closeness to real records.
Software
You do not need to code these from scratch
| Tool | What it offers |
|---|---|
| SDV (Python) | Gaussian copula, CTGAN, TVAE and CopulaGAN with a common interface, plus quality reports. |
| synthpop (R) | Sequential CART and parametric synthesis, widely used in health research and official statistics. |
| simstudy (R) | Simulating trial and study data from user-defined assumptions. |
| Synthea | Open-source simulator of synthetic patient records along care pathways. |
| Ecliptica® and others | Dedicated platforms for trial simulation and synthetic data, covered in the afternoon sessions. |
From real table to synthetic table in a few lines with SDV. Running the generator is the easy part; metadata, rules and evaluation are where the effort goes.
from sdv.metadata import Metadata
from sdv.single_table import GaussianCopulaSynthesizer
metadata = Metadata.detect_from_dataframe(data=real_df)
synth = GaussianCopulaSynthesizer(metadata)
synth.fit(real_df)
synthetic_df = synth.sample(num_rows=500)
# Try a GAN by changing one line:
# from sdv.single_table import CTGANSynthesizer
# synth = CTGANSynthesizer(metadata, epochs=300)Fidelity, utility and privacy
Evaluation
Three questions for any synthetic dataset
Fidelity
Does it look like the real data? Distributions, correlations, survival curves, clinical rules.
Utility
Does it give the same answers? The same analysis on real and synthetic data gives similar estimates.
Privacy
Could anyone be identified? No copies of real patients and no recognisable rare individuals.
These pull against each other: the closer the data are to the real data, the higher the privacy risk. The purpose decides the balance. See the trade-off in the lab →
Evaluation
Practical checks
Fidelity: look and compare
Side-by-side histograms per variable, correlation heatmaps, overlaid Kaplan–Meier curves, the share of records breaking clinical rules, and whether a classifier can tell real from synthetic.
Utility: redo the analysis
Run the same model on both, for example a Cox hazard ratio with 95% CI. Do the intervals overlap? Is the conclusion the same? For prediction models, train on synthetic and test on real (TSTR). Check subgroups too.
Privacy
Synthetic does not automatically mean anonymous
Measure exact and near copies (distance to the closest real record), test whether an attacker could tell if a person was in the training data (membership inference), and look for rare combinations that point to one person. Differential privacy adds mathematical guarantees, usually at some cost to accuracy.
Pitfalls
Five common mistakes
| Mistake | Instead |
|---|---|
| Choosing the most advanced method first | Define the purpose, then start simple. |
| Checking only averages and single histograms | Check relationships, subgroups and the main analysis. |
| Ignoring censoring, dates and clinical rules | Write the rules down and test them. |
| Assuming synthetic means anonymous | Measure privacy. |
| Not recording how the data were made | Document method, settings, checks and limitations. |