Synthetic data generator with validation

Functional — runs in your browser

Synthetic data generated in your browser. No real records.

What it shows

Generates, in your browser, a synthetic dataset for a real business process (SME property insurance underwriting or SME credit dossier review) and checks it immediately with seven controls computed on the generated data: schema, referential integrity, temporal coherence, target distributions, correlations, injected anomalies and the absence of personal data.

The demonstration

Synthetic data generator Ready
Domain pack
Engine (shared by all packs)
The same seed produces exactly the same dataset.
Number of rows
2 %
The anomaly kinds of the chosen pack, in equal parts; the list and the detection rules are under check 6.
Activity mix

  • generated
  • target distribution

Check results

Downloads

Exported files use the column names and values of the page language, and the JSON Schema describes exactly those names. The Romanian Excel variant opens straight into columns when Windows uses Romanian regional settings; the standard variant is for other software.

Check details

Preview: first 50 rows

How it works

The generator has a shared engine and domain packs: the engine holds the seeded pseudo-random generator (mulberry32), normal, lognormal, Poisson and categorical sampling, correlation through a Gaussian copula, export and the checks, while the pack describes the process (tables, rules, target distributions, anomaly types). This is how Virtual Soft works on projects too: one reusable engine and one domain pack per client.

In the insurance pack, the sum insured is lognormal, the number of claims per policy follows a Poisson distribution whose rate depends on construction, protections and zone, and the loss is lognormal, proportional to the sum insured and capped at it; only a loss above the deductible becomes a claim, which pays the loss minus the deductible. In the lending pack, firm age, turnover and indebtedness are drawn jointly through the Gaussian copula, and the decision follows a published rule set with a documented noise term.

Every observed statistic is compared with the value that follows analytically from the model, with a tolerance of 4.5 standard errors (for quantiles on the probability scale; for the loss ratio with claims capped at 250,000 RON), and when there are too few rows to confirm, the result is “not enough rows to confirm”, not “failed”; rows flagged by the anomaly rules are excluded from the statistics and reported separately. Anomaly rules are either exact, restating the defect and required to find every injected row with no false alarm, or heuristic (sum insured entered in thousands, the same policy entered twice under another ID, a premium off the tariff), which can miss small defects or flag clean rows; for these the page reports precision and recall and compares them with the values computed from the model.

Synthetic data is used in training and development for three reasons. The first is confidentiality: the dataset holds no personal data, so it falls outside the scope of the GDPR, and control stays with the client; learners' data in the training platform remains personal data and is processed under the Art. 28 agreement. The second is reproducibility: the same seed produces the same file, byte for byte. The third is the failure cases built deliberately into laboratories. Its limit is that it reproduces only the relationships that were modelled, not the unknown ones, and it does not replace validation on real, authorised data. Plausibility is checked against the stated target distributions, not against a real portfolio.

What is real and what is simulated

Real, computed in the page:

  • generating the rows, in your browser, with no data sent to the server;
  • all seven checks and every number in the validation panel;
  • the generation time shown;
  • the downloaded files (standard CSV and CSV for Romanian Excel, JSON, JSON Schema, HTML report, HTML technical sheet), with column names and values in the page language.

Simulated:

  • every record: policies, claims, dossiers and decisions are invented by the generator;
  • the model parameters (rates, frequencies, decision thresholds) are illustrative values chosen for plausibility, not taken from an insurer or a bank;
  • the risk zones Z1–Z5 and the activity groups are generic labels with no link to real places or firms.

Last verified: