Davies Meyer – home
    Data3 min read

    Synthetic Data

    Synthetic data is artificially generated information intended to reproduce particular structures, distributions or application scenarios. It can help develop, test or train systems. However, the label “synthetic” guarantees neither anonymity nor realism nor an appropriate lawful basis for the entire creation and use process.

    Synthetic Data explained

    A test system needs orders with different baskets, spellings or missing details. Such cases can be generated without copying actual customer records into every development environment. Depending on the task, data may come from rules, simulations or models that analyse real data. First define which properties the test needs and what claims it should be capable of supporting.

    Using real data for generation still processes that data. A subsequently generated dataset does not retrospectively resolve the earlier processing. Outputs may also reveal information about real people or inadequately represent properties of the source material. Assess provenance, permissible purpose and risks of creation and sharing separately. ICO guidance explains such technical limits; it does not provide legal clearance for deployment in Germany.

    Quality needs multiple checks. Similar individual distributions do not establish accurate relationships, rare cases or performance on a particular task. A model trained on synthetic data needs evaluation for its intended use against appropriate independent real data. Additional generated rows are not additional independent market observations. Synthetic clicks do not establish real conversion impact.

    Creative Engineering uses artificial data for specific purposes, such as checking unusual inputs and fragile workflows earlier. We take responsibility for the concept and quality. NIST distinguishes synthetic-data methods from correctly implemented differential privacy; they are not equivalent. Assess task quality, protection and full effort together. Where a simple artificial test case suffices, a complex generation model is not automatically better.

    Examples

    Hypothetical application

    A development team creates artificial orders with long addresses, missing optional fields and varied product combinations. These test form presentation and error handling. Whether the new design helps real people and enables more successful orders is investigated separately with suitable real usage data.

    Key Points

    • Artificially generated does not automatically mean anonymous or realistic.
    • Distinguish test data, training data and observed market data.
    • Assess task quality, protection and total effort separately.

    Practical application

    Define the test task and clearly label artificial data. Assess its usefulness and limits before deriving real-world decisions.

    Useful measures

    Task suitability

    Check whether the data covers intended test cases and relevant properties.

    Independent validation

    Evaluate later system performance using suitable data that does not merely repeat generation.

    Protection and full effort

    Assess disclosure risks, generation, review and ongoing maintenance together.

    Common mistakes

    • Treating the word synthetic as a privacy guarantee.
    • Counting generated records as new independent market observations.
    • Checking only average similarity while overlooking relevant edge cases.

    Sources and context

    Frequently Asked Questions about Synthetic Data

    That cannot be stated universally. Depending on generation, information about real people may be inferable; creating the data may also process personal source records.

    Loading related terms…

    All Terms