Fully synthetic
Every record is generated. Strongest privacy, but relies most on the quality of the generator.
One-day course for clinical, regulatory and HTA professionals
…without being anyone. Synthetic data can speed up trial design, open up access to evidence and protect privacy. This companion site walks through how it is made, how to check it, and how to choose a method you can defend.
Try it before the theory
What synthetic data is, and is not
Artificially generated data that reproduces the patterns and statistical properties of real data, without being a record of any real person.
Every record is generated. Strongest privacy, but relies most on the quality of the generator.
Only the sensitive variables are replaced. The rest of the record stays real.
Generated from assumptions or models, such as event rates or PK/PD equations, without learning from patient data.
Start with one question
The purpose decides how realistic the data must be, which method to use and how you must check it. The bars show how much realism each purpose needs.
Build and test code and pipelines without touching patient data.
Explore trial designs, sample sizes and what-if scenarios before enrolment.
Reproduce the real analysis, such as a hazard ratio, so relationships must be preserved.
Give collaborators data that behaves like the original while protecting individuals.
Interactive simulation lab
Two experiments run in your browser: simulate a whole trial from assumptions, then compare three ways of generating synthetic patients.
Switch between independent columns, a copula and a bootstrap, and watch fidelity rise while the share of near-copies of real patients climbs with it.
Run the experiment →Assume → Simulate → Compare → Decide
Change one setting at a time and read what responds.
Four families of methods
Move to a more flexible family only when the simpler method is not good enough for your purpose.
Build data from assumptions and known models.
Read more →Reuse and perturb real records: bootstrap, noise, SMOTE.
Read more →Learn distributions and correlations: copulas, Bayesian networks, CART.
Read more →Learn complex patterns: GANs, VAEs, diffusion, transformers.
Read more →The programme
Timings are approximate and may shift on the day.
10:45 to 12:00 with Dr Saqib Ur Rehman. Everything in the session, plus simulations, a method chooser and practice questions, is on this site.
Go to the session contentMethod chooser
Answer four short questions about your purpose and your data. You get a starting method, alternatives, tools and the checks you will need.