Session 2 · 10:45 to 12:00 · three parts of about 25 minutes

Methods for generating synthetic data

From simple simulation to deep learning: what each method does, when to use it, and what to watch out for. No equations, just the ideas you need to choose well.

01The methods toolbox

Simple, transparent and often enough

These methods should be tried first. For trial-sized datasets they are frequently as good as anything more complex.

Family 01 · Rules and simulation

Knowledge-driven simulation

You write down what you believe about the population and the trial, then let the computer generate patients. Statisticians have done this for decades to check sample sizes and trial designs. Examples include a two-arm trial with assumed event rates and dropout, PK/PD models linking dose to concentration and response, and Synthea, which simulates patient records along care pathways.

Use whenDesigning studies, estimating power, exploring what-if scenarios, or when no real data exist yet.
Watch outThe data are only as good as the assumptions. Simulation will not reveal anything you did not put in.

A two-arm survival trial in about ten lines of Python. Every number is an assumption you can change. Try it interactively in the lab →

Python
import numpy as np, pandas as pd
rng = np.random.default_rng(2026)

n = 300                              # patients per arm
median_ctrl, hr = 6.0, 0.75          # assumptions (months, hazard ratio)
lam_c = np.log(2) / median_ctrl      # exponential hazard, control
arm   = np.repeat(["Placebo", "Active"], n)
lam   = np.where(arm == "Active", lam_c * hr, lam_c)
t_evt = rng.exponential(1 / lam)     # time to event
t_cen = rng.uniform(6, 24, 2 * n)    # administrative censoring
df = pd.DataFrame({"arm": arm,
     "time":  np.minimum(t_evt, t_cen),
     "event": (t_evt <= t_cen).astype(int)})

Family 02 · Resampling

Resampling and perturbation

Bootstrap draws real records with replacement. Noise addition makes small random changes so no record is an exact copy. SMOTE-type methods create new records between existing ones, often to boost rare classes.

Quick and easy, but records stay close to real patients. Good for internal testing and class balancing. Rarely suitable for sharing outside the organisation without privacy checks. See the near-copy risk in the lab →

Family 03 · Statistical models

The Gaussian copula in three steps

  1. Learn each column. What does age look like on its own? Tumour size? Learn the shape of every variable.
  2. Learn how they move together. Older patients have more comorbidities. Learn these correlations separately.
  3. Generate new rows. Draw correlated random numbers and map them back onto each column's shape.
StrengthsFast, stable with small samples, easy to explain to clinicians and regulators. The best first baseline for tabular data.
LimitsCaptures mainly pairwise, smooth relationships. Can miss interactions and hard clinical rules.

Family 03 · Statistical models

Build a patient one variable at a time

Bayesian networks and sequential CART (for example the R package synthpop) generate each variable from the ones already generated, much like following a clinical story.

Age→Sex→Disease stage→Treatment→Outcome

It handles mixed data types and nonlinear relationships well, lets you build clinical logic into the variable order, and works with a few hundred to a few thousand records. The order matters, and very many variables make it slow.

R
library(synthpop)
syn_obj <- syn(real_df, method = "cart", seed = 2026)
synthetic_df <- syn_obj$syn
compare(syn_obj, real_df)   # side-by-side distributions
02Deep learning generators

When relationships get complex

Family 04 · Deep learning

Why deep learning?

Deep generators help when there are complex, interacting relationships, hundreds of variables, sequences over time such as visits and treatment histories, or unstructured text and images.

The priceMore data, more tuning, more computing, and harder to explain. With only a few hundred trial patients, deep learning does not automatically win.

Family 04 · Deep learning

GANs: a forger and a detective

A generator (the forger) turns random noise into fake patients. A discriminator (the detective) tries to tell fake from real. Each round of feedback makes both better, and training stops when the detective can no longer tell the difference.

Tabular versions for mixed clinical data include CTGAN, CopulaGAN and Wasserstein GANs such as WGAN-GP and conditional tabular WGANs.

Watch outTraining can be unstable. “Mode collapse” means the generator only learns common patient types and misses rare subgroups, which are often the ones that matter most.

Family 04 · Deep learning

Three more families worth knowing

VAE

VAEs (e.g. TVAE)

Compress each patient into a small summary, then learn to rebuild patients from it. Often more stable than GANs.

DDPM

Diffusion (e.g. TabDDPM)

Add noise step by step, then learn to remove it. Start from noise to get new patients. High quality, slower.

LLM

Transformers and LLMs

Treat a patient history as a sentence of events and generate the next event. Promising for EHR sequences; privacy needs care.

Clinical reality

What makes clinical data hard to synthesise

ChallengeWhat to check
Survival and censoringTime and event status stay consistent; Kaplan–Meier curves match.
Small samplesTrials often have hundreds, not millions, of patients. Prefer stable methods.
Rare eventsSerious adverse events and small subgroups are still present.
Clinical rulesDeath after randomisation, doses within range, valid lab units.
Mixed data typesContinuous, categorical, dates and text in one dataset.
Longitudinal structureRepeated visits per patient remain coherent over time.

Case study

Synthetic data from a lung cancer trial

Using a randomised lung cancer trial (Tarceva versus placebo) with overall survival and progression-free survival endpoints, provided by RSS under a data-sharing agreement, we compared a Gaussian Copula Synthesizer with a conditional tabular Wasserstein GAN.

Both were judged with the same checks: Kaplan–Meier curves, hazard ratios, distributions and closeness to real records.

Results to be addedKaplan–Meier overlays and the hazard-ratio comparison will be added here once cleared for external use.

Software

You do not need to code these from scratch

ToolWhat it offers
SDV (Python)Gaussian copula, CTGAN, TVAE and CopulaGAN with a common interface, plus quality reports.
synthpop (R)Sequential CART and parametric synthesis, widely used in health research and official statistics.
simstudy (R)Simulating trial and study data from user-defined assumptions.
SyntheaOpen-source simulator of synthetic patient records along care pathways.
Ecliptica® and othersDedicated platforms for trial simulation and synthetic data, covered in the afternoon sessions.

From real table to synthetic table in a few lines with SDV. Running the generator is the easy part; metadata, rules and evaluation are where the effort goes.

Python · SDV
from sdv.metadata import Metadata
from sdv.single_table import GaussianCopulaSynthesizer

metadata = Metadata.detect_from_dataframe(data=real_df)
synth = GaussianCopulaSynthesizer(metadata)
synth.fit(real_df)
synthetic_df = synth.sample(num_rows=500)

# Try a GAN by changing one line:
# from sdv.single_table import CTGANSynthesizer
# synth = CTGANSynthesizer(metadata, epochs=300)
03Is it good enough?

Fidelity, utility and privacy

Evaluation

Three questions for any synthetic dataset

01

Fidelity

Does it look like the real data? Distributions, correlations, survival curves, clinical rules.

02

Utility

Does it give the same answers? The same analysis on real and synthetic data gives similar estimates.

03

Privacy

Could anyone be identified? No copies of real patients and no recognisable rare individuals.

These pull against each other: the closer the data are to the real data, the higher the privacy risk. The purpose decides the balance. See the trade-off in the lab →

Evaluation

Practical checks

Fidelity: look and compare

Side-by-side histograms per variable, correlation heatmaps, overlaid Kaplan–Meier curves, the share of records breaking clinical rules, and whether a classifier can tell real from synthetic.

Utility: redo the analysis

Run the same model on both, for example a Cox hazard ratio with 95% CI. Do the intervals overlap? Is the conclusion the same? For prediction models, train on synthetic and test on real (TSTR). Check subgroups too.

Privacy

Synthetic does not automatically mean anonymous

Measure exact and near copies (distance to the closest real record), test whether an attacker could tell if a person was in the training data (membership inference), and look for rare combinations that point to one person. Differential privacy adds mathematical guarantees, usually at some cost to accuracy.

Deep generators can memorise training records, especially with small data and long training. Always test.

Pitfalls

Five common mistakes

MistakeInstead
Choosing the most advanced method firstDefine the purpose, then start simple.
Checking only averages and single histogramsCheck relationships, subgroups and the main analysis.
Ignoring censoring, dates and clinical rulesWrite the rules down and test them.
Assuming synthetic means anonymousMeasure privacy.
Not recording how the data were madeDocument method, settings, checks and limitations.