Synthetic data generator with validation
Synthetic data generated in your browser. No real records.
What it shows
Generates, in your browser, a synthetic dataset for a real business process (SME property insurance underwriting or SME credit dossier review) and checks it immediately with seven controls computed on the generated data: schema, referential integrity, temporal coherence, target distributions, correlations, injected anomalies and the absence of personal data.
The demonstration
- generated
- target distribution
Check results
Downloads
Exported files use the column names and values of the page language, and the JSON Schema describes exactly those names. The Romanian Excel variant opens straight into columns when Windows uses Romanian regional settings; the standard variant is for other software.
Check details
Preview: first 50 rows
How it works
The generator has a shared engine and domain packs: the engine holds the seeded pseudo-random generator (mulberry32), normal, lognormal, Poisson and categorical sampling, correlation through a Gaussian copula, export and the checks, while the pack describes the process (tables, rules, target distributions, anomaly types). This is how Virtual Soft works on projects too: one reusable engine and one domain pack per client.
In the insurance pack, the sum insured is lognormal, the number of claims per policy follows a Poisson distribution whose rate depends on construction, protections and zone, and the loss is lognormal, proportional to the sum insured and capped at it; only a loss above the deductible becomes a claim, which pays the loss minus the deductible. In the lending pack, firm age, turnover and indebtedness are drawn jointly through the Gaussian copula, and the decision follows a published rule set with a documented noise term.
Every observed statistic is compared with the value that follows analytically from the model, with a tolerance of 4.5 standard errors (for quantiles on the probability scale; for the loss ratio with claims capped at 250,000 RON), and when there are too few rows to confirm, the result is “not enough rows to confirm”, not “failed”; rows flagged by the anomaly rules are excluded from the statistics and reported separately. Anomaly rules are either exact, restating the defect and required to find every injected row with no false alarm, or heuristic (sum insured entered in thousands, the same policy entered twice under another ID, a premium off the tariff), which can miss small defects or flag clean rows; for these the page reports precision and recall and compares them with the values computed from the model.
Synthetic data is used in training and development for three reasons. The first is confidentiality: the dataset holds no personal data, so it falls outside the scope of the GDPR, and control stays with the client; learners' data in the training platform remains personal data and is processed under the Art. 28 agreement. The second is reproducibility: the same seed produces the same file, byte for byte. The third is the failure cases built deliberately into laboratories. Its limit is that it reproduces only the relationships that were modelled, not the unknown ones, and it does not replace validation on real, authorised data. Plausibility is checked against the stated target distributions, not against a real portfolio.
What is real and what is simulated
Real, computed in the page:
- generating the rows, in your browser, with no data sent to the server;
- all seven checks and every number in the validation panel;
- the generation time shown;
- the downloaded files (standard CSV and CSV for Romanian Excel, JSON, JSON Schema, HTML report, HTML technical sheet), with column names and values in the page language.
Simulated:
- every record: policies, claims, dossiers and decisions are invented by the generator;
- the model parameters (rates, frequencies, decision thresholds) are illustrative values chosen for plausibility, not taken from an insurer or a bank;
- the risk zones Z1–Z5 and the activity groups are generic labels with no link to real places or firms.
Last verified: