What is the Synthetic Data Module?

The Synthetic Data Module generates privacy-safe "digital twins" of real datasets. You provide an Excel source dataset, and the module produces a brand-new synthetic dataset that preserves the original's statistical structure -- distributions, correlations, and effect sizes -- without containing a single one of your original data points.

The result is a dataset you can share, publish, and analyse freely -- for collaboration, teaching, machine-learning training, or methods development -- without exposing confidential or regulated information.


What "synthetic data" means here

Synthetic data mirrors the statistical patterns of a real dataset without reusing any actual records. Because no original observation survives into the output, the synthetic dataset can be distributed without the privacy, consent, and regulatory constraints that apply to the source data. The statistical relationships -- how variables are distributed, how they correlate, how large the effects between them are -- are reconstructed so that analyses reach the same conclusions they would on the original.


The four-step workflow

StepWhat happens
1. Upload Your DataProvide a raw dataset as an Excel .xlsx or .xls file
2. Automatic AnalysisThe module detects each variable's type, distribution, and relationships
3. Configure SettingsSet target sample size, review correlations and effect sizes, fine-tune parameters
4. Generate & ValidateThe engine produces the synthetic dataset and a full validation report in seconds

What the module preserves

CapabilityWhat it does
Privacy ProtectionSynthetic dataset can be shared without exposing any original data point
Statistical IntegrityCorrelation structures preserved using Cholesky decomposition
Effect Size MatchingCohen's d, Odds Ratios, and Hazard Ratios are maintained
Distribution FittingVariable distributions (Normal, Uniform, Binomial, etc.) detected and replicated
Visual ValidationLove plots, correlation matrices, and SMD comparisons produced every run
Fast ProcessingOptimised algorithms generate well over 100,000 rows in seconds

Who the module is for

AudienceTypical uses
Clinical ResearchPower-analysis calculations, method development, IRB-compliant datasets, multi-site studies
Machine LearningTraining-data augmentation, class-imbalance handling, model validation, feature engineering
EducationStatistics coursework, workshop demonstrations, student exercises, publication examples

How accurate is the generated data?

The module preserves the large majority of a dataset's statistical relationships, and every generated dataset is accompanied by a validation report showing standardised-mean-difference (SMD) analysis and quality scores, so you can verify the fidelity of each run.


Supported data

The module supports six variable types -- continuous, binary, categorical, count, time-to-event, and proportion -- and detects each type automatically. Datasets can range from a few hundred rows to well over 100,000 rows in a single run.


Data sources

The current workspace accepts a raw dataset in Excel .xlsx or .xls format. It does not currently expose summary-statistics, HTML-table, or manual-template input routes in the interface.


Privacy and data handling

The module is built privacy-first: uploads are encrypted, and the original data is used only to analyse its statistical structure -- it is not retained after processing. Because the synthetic output contains no original records, it sidesteps the confidentiality and regulatory constraints that govern the source data.


Access states

StateWhat you see
Not signed inRedirected to the sign-in page with a return link
Signed in, without All AccessThe module remains locked; choose an All Access plan from the dashboard or sidebar
Signed in, with All AccessFull generation workflow