AI 日报hiw3c.com

数据模仿-不要让您的编码代理发明自己的测试世界

原文标题 · Datamimic – don't let your coding agent invent its own test world
Hacker News Top github.com 网页快照
正文为英文,可一键机器翻译(仅首次需要等待)

Uh oh!

There was an error while loading. Please reload this page .

Notifications You must be signed in to change notification settings

Latest commit

History

Folders and files

Repository files navigation

DATAMIMIC — Governed Test Data for Regulated Enterprises

This repository contains the DATAMIMIC Community Edition (CE). MIT-licensed, Python-native, MCP-ready. CE is fully usable standalone for deterministic synthetic data generation and PII-aware pseudonymization. The Enterprise Platform adds governed workflows, PII scanning, role-based access, audit logging, scheduling, multi-system execution, and the full operational layer that regulated enterprises require. 👉 Enterprise Platform: datamimic.io | 📘 Docs: docs.datamimic.io | 📅 Book a strategy call: datamimic.io/contact 🤖 AI agent? Start at AGENTS.md and use the project CLI: preserve new intent as model.dm.json , submit an early best attempt via datamimic scaffold ... --format json , repair from the structured issues, declare an expectation per stated requirement, and stop on verified=true . Existing raw XML uses lint plus bounded dry-run.

What is DATAMIMIC?

DATAMIMIC CE is the open-source deterministic data engine at the core of the DATAMIMIC Enterprise Platform. It is usable standalone for synthetic data generation and PII-aware pseudonymization in any local, CI, or agent-driven workflow.

The Enterprise Platform adds the governed workflows, scanners, dashboards, and execution layer that regulated enterprises require for production-scale test-data operations.

Generate fully synthetic, deterministic datasets — model-driven, no source data required

Pseudonymize staging/QA exports — deterministic (seeded) or privacy-maximized (non-seeded) field transformation; PII fields identified and modeled manually in the XML pipeline

Execute single-system pipelines against PostgreSQL · MySQL · Oracle · MS SQL · SQLite · MongoDB · CSV · JSON · XML · XLSX · DbUnit · fixed-width ( .fcw )

Model behavior — weighted state machines, composite multi-field references, control flow ( <while> , <assert> ), and a scriptable memstore for staged aggregation

Emit provenance — append-only execution logs and per-output content hash for audit re-execution

Guide agents — machine-readable capabilities, progressive reference queries, and one canonical CLI scaffold transaction; an optional MCP adapter exposes the same authoring service

PII scanner — probability-scored field detection with configurable thresholds via DataWorkbench

Multi-system execution — Oracle / MongoDB / Kafka in coordinated workflows with referential integrity

Industry message templates — EDIFACT / SWIFT MT / HL7 v2.x / HL7 FHIR generated as deterministic test/training artefacts

Governance layer — role-based dashboards, audit trails, approval flows, reusable enterprise templates, scheduler

Performance core — Rust fastpath, ML/auto-regressive engine for complex distributions, keyset and manifest building, optimised distributed execution

On-premise / air-gapped deployment — podman-compose or Helm, with consulting-led rollout

Deployed in regulated EU banking environments for deterministic test data across Oracle, MongoDB, and Kafka pipelines. Reference customers available under NDA — see also datamimic.io case studies .

AI agents: author, verify, and run data models

The CLI is the baseline agent contract. Install CE with pip install datamimic-ce ; inside this checkout, use .venv/bin/datamimic so a stale global installation cannot change the available schema or commands.

capabilities , authoring-reference projections, and the commands shown with --format json return machine-readable JSON. On a failed scaffold attempt, change model.dm.json using its structured validation issues, typed repair, or rule diagnostics before retrying. A typed max_count remediation instead changes only the bounded scaffold parameter to at least its reported minimum. Never repeat an identical failed call. A successful scaffold result is terminal for authoring, so do not lint or dry-run its generated XML again. Exact source fragments are discoverable through queries such as --category source --kind memstore .

Optional MCP adapter

When the calling environment already exposes DATAMIMIC MCP tools, they map to the same canonical contracts and implementations: reference → datamimic_reference , scaffold → datamimic_scaffold , lint → datamimic_check , and dry-run → datamimic_run . Install the adapter with pip install "datamimic-ce[mcp]" ; registration details belong in the MCP quickstart , not in the authoring workflow. The adapter intentionally exposes only the four canonical reference, scaffold, check, and bounded-run operations; domain generation remains a Python/CLI capability rather than a parallel MCP authoring path.

Prompts to paste into your agent

Create the dataset I describe with DATAMIMIC. Read AGENTS.md first. In a repository checkout use `.venv/bin/datamimic`; otherwise use the current `datamimic` CLI. Preserve my intent as `model.dm.json`; do not hand-write XML. Start from the minimal valid document shape in AGENTS.md ("Authoring a new model"). Two rules prevent most rejections: the top level allows ONLY version, seed, products, expectations; product-level "kind" (generated/source/time_series) is a different vocabulary from field-level "kind" (increment, values, weighted, int_range, decimal_range, pattern, constant, script). Range fields take minimum/maximum, never min/max. Submit EARLY: run `datamimic scaffold model.dm.json --format json` with your best attempt after at most one discovery call. Repair from the structured issues (path/code/message/allowed_fields) and diagnostics (fix_hint) — they teach the schema faster than more discovery. Never resubmit an unchanged document. If a remediation requests a larger max_count, retry scaffold with at least that value without changing the intent. Declare an expectation for every requirement I state (counts as exact_count with a "count" field, uniqueness, allowed values, ranges, foreign keys) — verified=true certifies only what you declared. Stop on verified=true; do not lint or dry-run the generated XML. If I request real execution, save the returned XML as a generated artifact and run that descriptor. Return the model.dm.json path and concise verification evidence.

Relational hierarchy with referential integrity (fully supported — no XML needed)

Seed a relational dataset with referential integrity: 4 customers, each with exactly 2 orders. Customers get an incrementing unique id and a region from {north, south, east, west}. Each order carries the REAL parent customer id as a foreign key and an amount between 10.0 and 500.0. Follow AGENTS.md's "Authoring a new model" and its structural recipes: orders nest inside the customer product's "children" array; the FK field is {"kind": "script", "script": "parent.id"} with a foreign_key role — a randomly generated FK passes schema validation but fails per-parent-count acceptance. Declare expectations for the customer count, customer id uniqueness, exactly 2 orders per customer (per_parent_count), the orders->customers foreign key, and the amount range. Stop on verified=true and show the acceptance evidence.

Raw XML remains supported for existing descriptors (lint → dry-run → run; see AGENTS.md). For new models it is a last resort: only when a scaffold issue explicitly classifies the requirement as unsupported_intent should an agent hand-author XML, preserving that evidence.

CE vs Enterprise Platform

CE and EE are not the same engine with a feature flag . They share the DSL and determinism contract, but EE is an independently optimised execution engine built for enterprise-scale throughput and operational control.

Engine comparison

Platform capabilities (EE only)

👉 Explore the Enterprise Platform | Book a platform demo

EE runtime profiles

The EE core supports three runtime configuration profiles, selectable per execution context:

Logging depth is independently configurable per profile — from minimal (throughput-optimised) to full nested tracing across importers, exporters, and generation stages.

EE template engine

The EE template engine generates industry-standard financial messages from DATAMIMIC models. The workbench parses uploaded message samples, auto-detects the message type, and validates edits against the registered spec version in real time.

Capabilities

Spec-aware form editing — segments and elements rendered as structured forms with mandatory/optional indicators, per-field value suggestions, and inline custom-extension support

Strict validation against baked spec versions, with segment- and element-level error reporting

Advisory mode when a spec is unregistered or in draft — editing stays enabled, validation continues as guidance

Round-trip between the structured form view and the authoritative template text — no fidelity loss

Download / adjust / upload your own spec — customers can extend or override the baked spec catalogue without waiting for a release

Live structure tree + preview for every edit

File auto-detection — upload an existing message, the editor identifies the type and loads the matching spec

Format coverage

Customers can extend the spec catalogue between releases by downloading, adjusting, and uploading their own spec files directly.

Generated messages are deterministic and traceable to their source model, and syntactically valid against the registered spec. They are intended for test and training environments only — they are not network-validated and must not be transmitted on production SWIFTNet or EDI networks. See the SWIFT CSP note below.

Who is DATAMIMIC for?

Enterprise Platform (EE)

Community Edition (CE)

Developers and data engineers who need deterministic synthetic data generation or PII-aware pseudonymization in local environments, CI pipelines, or agent-driven workflows. PII field identification is manual — the EE DataWorkbench automates this step.

Why deterministic generation matters

Most test data tools produce random output. That breaks regression tests, audit trails, and cross-team reproducibility.

Same engine version + same model + same seed = byte-identical output , every run, every machine. Holds at three layers: the generate_domain facade, every domain service called directly, and every literal generator that accepts an rng= argument. Verified per-service on every CI run via tests_ce/architecture/test_service_replay_determinism.py .

DSL-level seeding: <setup rngSeed="N"> makes the whole model deterministic — every seed-less <variable entity="…"> derives a reproducible child RNG from it, and <variable rngSeed="…"> overrides it for that block (no seed anywhere → wall-clock random). Verified by tests_ce/integration_tests/test_determinism_seed_scenarios . As of 4.0.0 the same seed also reaches standalone literal generators ( <key generator="…"> ), typed/pattern keys, DateTimeGenerator , and cross-page unique picks — machine-independently.

Source reads: distribution="ordered" reads a data source in stable file order; distribution="random" shuffles but replays identically when <setup rngSeed> is set (without a seed the shuffle is non-deterministic by design, for privacy-maximized one-time deliveries). Deterministic shuffling across distributed / multi-process execution is EE.

Provenance hash on every facade output = re-executable lineage. Same input → same determinism_proof.content_hash , always.

UUIDv5 entity identifiers = stable across runs and machines.

Single wall-clock SPOT ( now_utc_naive() ); raw datetime.now() is forbidden in production code and the clock-drift architecture gate fails CI on any reintroduction.

RNG/clock runtime SPOTs in datamimic_ce/domains/domain_core/runtime/ : spawn_rng (reproducible child-RNG derivation), now_utc_naive , and resolve_clock . The same contract vocabulary the Enterprise Platform enforces end-to-end.

The Enterprise Platform (EE) goes further: beyond the CE contract, EE makes the whole execution environment deterministic — a configurable/frozen wall-clock (not just CE's fixed anchor), and deterministic SAFE_GLOBALS plus the Python random functions, so sandboxed script expressions and any stdlib random call replay identically as well.

from datamimic_ce . domains . facade import generate_domain request = { "domain" : "person" , "version" : "v1" , "count" : 1 , "seed" : "regression-suite-42" , # identical seed → identical output "locale" : "en_US" , "clock" : "2025-01-01T00:00:00Z" # fixed clock = stable time context } response = generate_domain ( request ) # response["determinism_proof"]["content_hash"] is stable across runs.

Direct service use is equally deterministic when given a seeded RNG:

import random from datamimic_ce . domains . finance . services import CreditCardService # Same seeded Random → byte-identical CreditCard across runs. card_a = CreditCardService ( rng = random . Random ( 42 )). generate () card_b = CreditCardService ( rng = random . Random ( 42 )). generate () assert card_a . bic == card_b . bic and card_a . card_number == card_b . card_number

Determinism contract — CE vs EE

CE delivers contract-enforced determinism for the synthetic-data generation surface (facade, services, generators) and, as of 4.0.0, for seeded XML descriptors — byte-identical across machines, executed single-process. The Enterprise Platform extends the same contract to distributed and multi-system execution with referential integrity and the seeded/unseeded pseudonymization modes, and adds the five drift-gates that lock the contract end-to-end for regulated deployments.

How DATAMIMIC differs from Faker and generic generators

# Faker — broken relationships from faker import Faker fake = Faker () patient_age = fake . random_int ( 1 , 99 ) conditions = [ fake . word ()] # "25-year-old with Alzheimer's" — meaningless for any real test # DATAMIMIC — domain-aware, deterministic with a seed import random from datamimic_ce . domains . healthcare . services import PatientService patient = PatientService ( rng = random . Random ( 42 )). generate () print ( f" { patient . full_name } , { patient . age } , { patient . conditions } " ) # Age-appropriate, domain-consistent — and identical every run with a fixed seed

Quickstart — Community Edition

pip install datamimic-ce

Healthcare domain

import random from datamimic_ce . domains . healthcare . services import PatientService patient = PatientService ( rng = random . Random ( 42 )). generate () print ( patient . full_name , patient . age , patient . conditions ) # Age-appropriate conditions, demographically realistic; deterministic with a seed

Finance domain

import random from datamimic_ce . domains . finance . services import BankAccountService account = BankAccountService ( rng = random . Random ( 42 )). generate () print ( account . account_number , account . balance ) # Balance-consistent, locale-correct; reproducible with a seed

Pseudonymization — CE (manual model)

DATAMIMIC supports two pseudonymization modes with different privacy postures:

Note on GDPR anonymization: Full anonymization status under GDPR depends on complete field coverage across all quasi-identifiers and a re-identification risk assessment on the complete record — not on individual field transformation alone. DATAMIMIC does not make anonymization claims on behalf of the customer. Non-seeded mode maximizes privacy at the transformation level; the customer is responsible for assessing re-identification risk across the full dataset.

In CE, PII fields are identified and modeled manually in the XML pipeline:

< setup defaultSeparator = " , " > < generate name = " customers " source = " customer_export.csv " target = " CSV " distribution = " ordered " > <!-- distribution="ordered" reads the source in a stable order — required so the Nth source row maps to the same seeded synthetic value on every run. The default ("random") shuffles non-deterministically and would break it. rngSeed on the <variable> makes the synthetic values reproducible; drop rngSeed for the privacy-maximized (non-deterministic) mode. --> < variable name = " p " entity = " Person " dataset = " DE " rngSeed = " 42 " /> < variable name = " acc " entity = " BankAccount " dataset = " DE " rngSeed = " 42 " /> < key name = " first_name " script = " p.given_name " /> < key name = " last_name " script = " p.family_name " /> < key name = " email " script = " p.email " /> < key name = " iban " script = " acc.iban " /> < key name = " birth_date " script = " p.birthdate " /> </ generate > </ setup >

Built-in converters can additionally transform a key's value — e.g. irreversibly hash the original instead of replacing it, or partially mask it:

< key name = " email " script = " p.email " converter = " Hash('sha256','hex') " /> < key name = " iban " script = " acc.iban " converter = " MiddleMask(8, 4) " />

Available converters (13): Mask , MiddleMask(start, end) , CutLength(n) , Substring(start, end) , JavaHash , RemoveNoneOrEmptyElement , Hash(type, format[, salt]) , DateFormat(fmt) , Append , UpperCase , LowerCase , Date2Timestamp , Timestamp2Date .

datamimic run ./pseudonymize-customers/datamimic.xml

source is a controlled export or staging input — never a live production connection.

With rngSeed set: same source record → same pseudonymized output on every run. Stable for regression testing.

Without rngSeed : non-deterministic output — no reversible mapping exists at the field level. Stronger privacy posture for one-time delivery scenarios.

In the Enterprise Platform (EE): the DataWorkbench PII scanner automatically scans source schemas, assigns probability scores to each field, and flags candidates above a configurable threshold. Flagged fields are wired into the pseudonymization model automatically — no manual field mapping required.
< setup > < generate name = " patients " count = " 1000 " target = " CSV " > < variable name = " patient " entity = " Patient " dataset = " US " ageMin = " 60 " ageMax = " 80 " rngSeed = " 42 " /> < key name = " full_name " script = " patient.full_name " /> < key name = " age " script = " patient.age " /> < array name = " conditions " script = " patient.conditions " /> </ generate > </ setup &