One-day course for clinical, regulatory and HTA professionals

Data that behaves like patients.

…without being anyone. Synthetic data can speed up trial design, open up access to evidence and protect privacy. This companion site walks through how it is made, how to check it, and how to choose a method you can defend.

Session 2: Methods for generating synthetic data Dr Saqib Ur Rehman. Three parts of about 25 minutes: the methods toolbox, deep learning generators, and how to judge whether the data are good enough. Read the session content →
4method families
2live simulations
15practice questions

Try it before the theory

Which points are real?

PurposeMethodEvaluateDocument

What synthetic data is, and is not

Patterns of real data, records of no one

Artificially generated data that reproduces the patterns and statistical properties of real data, without being a record of any real person.

1

Fully synthetic

Every record is generated. Strongest privacy, but relies most on the quality of the generator.

2

Partially synthetic

Only the sensitive variables are replaced. The rest of the record stays real.

3

Simulated

Generated from assumptions or models, such as event rates or PK/PD equations, without learning from patient data.

Three common misconceptions

“It is fake, so it is useless.”
It is useful when it reproduces the patterns a task needs, and that can be tested.
“It is always anonymous.”
Privacy must be measured. Poorly generated data can contain near-copies of real patients.
“It will replace clinical trials.”
It complements trials by supporting design, comparison and extrapolation.
02Purpose first

Start with one question

What must the data do?

The purpose decides how realistic the data must be, which method to use and how you must check it. The bars show how much realism each purpose needs.

Purpose 01

Test and train

Build and test code and pipelines without touching patient data.

Realism needed: low
Purpose 02

Design and simulate

Explore trial designs, sample sizes and what-if scenarios before enrolment.

Realism needed: medium
Purpose 03

Answer questions

Reproduce the real analysis, such as a hazard ratio, so relationships must be preserved.

Realism needed: high
Purpose 04

Share safely

Give collaborators data that behaves like the original while protecting individuals.

Realism needed: high + privacy checks
03Experiment room

Interactive simulation lab

Test an idea before trusting a method

Two experiments run in your browser: simulate a whole trial from assumptions, then compare three ways of generating synthetic patients.

04Method families

Four families of methods

Start on the left. Move right only when you must.

Move to a more flexible family only when the simpler method is not good enough for your purpose.

05Course day

The programme

From fundamentals to future directions in one day

Timings are approximate and may shift on the day.

This site covers

Session 2: Methods for generating synthetic data

10:45 to 12:00 with Dr Saqib Ur Rehman. Everything in the session, plus simulations, a method chooser and practice questions, is on this site.

Go to the session content
  1. Registration and networking
  2. Break
  3. Lunch
  4. Regulatory Perspectives on Synthetic DataMHRA perspective, with a response from RSS
  5. Applications of Synthetic Data: Case Studies using EclipticaIntroduction to Ecliptica, download and install; using Ecliptica for clinical trial simulation & PK/PD; external control arms
  6. Break
  7. Applications of Synthetic DataReal World Data: target trial emulation & digital twins example; Health Technology Assessment for JCA
06Decision corner

Method chooser

Not sure which method fits your project?

Answer four short questions about your purpose and your data. You get a starting method, alternatives, tools and the checks you will need.

A starting methodWhat else to tryThe checks you need
Open the method chooser