Test Data Compliance Boundary: Why Copying Real User Data Into Test Environments Is Already Illegal
A common routine that is already unlawful
"Pull a copy of the production database into staging" is daily routine for countless teams — and one of the most overlooked violations under China's Personal Information Protection Law. PIPL Article 6 establishes purpose limitation: personal information may only be processed for the purpose the user consented to. Consent for "providing the service" does not extend to testing — a new processing purpose requires a new legal basis, and the same conclusion holds under GDPR.
The practical risk compounds it: test environments consistently run weaker security than production (weak passwords, broader internal access, full logging, a database dump on every developer's laptop). Copying real data there means duplicating a protected asset into an undefended zone, and post-incident reviews of real breaches routinely find test databases involved.
So the right question is not "how do we move production data into testing safely" — it is "how do we make testing not need real data at all."
Four legitimate sources of substitute data
| Source | Fits when | Risk to watch |
|---|---|---|
| Pseudonymized / masked production data | You need real data distribution (performance, bug reproduction) | A hashed phone number is still reversible via rainbow tables — under PIPL Article 73, only true anonymization ("irreversible") escapes the personal-information definition; pseudonymized data does not |
| Public test ranges | Payment / SMS gateway integration | The gateway-published test card numbers (Visa 411111 and friends) fail in production by design — they are officially sanctioned ammunition |
| Synthetic data | Load tests, demos, teaching | Must be "format-valid yet random" — the design core of any decent generator |
| Hand-made data | Feature verification | Hand-crafted samples are often format-invalid, which means your validation logic never gets tested |
Synthetic data is the only source that combines format realism with zero individual linkage — provided the generator gets the checksums right.
How "format-valid but random" is actually done
Chinese national-format numbers are structured, not random digits. Take the resident ID (GB 11643):
- First 6 digits — administrative division code: the first two are a fixed province table (11 Beijing, 31 Shanghai, 44 Guangdong…); a 99 opening gives it away instantly
- 8-digit birthdate — must be a real calendar date (February 30 fails)
- 3-digit sequence — odd for male, even for female; this digit carries gender semantics
- Final check digit — the first 17 digits take weights 7,3,1,7,3,1…, summed mod 11, mapped onto 1 0 X 9 8 7 6 5 4 3 2 (ISO 7064 MOD 11-2). Flip any digit and the check fails
The Unified Social Credit Code (GB 32100-2015) uses a 31-character alphabet (I/O/S/V/Z deliberately excluded against confusion) with a mod-31 check; bank cards run Luhn; phone numbers carry no check digit but the prefix table is a hard constraint (the 12x prefix is retired).
The acceptance test for synthetic data: it passes a standards-based validator (our resident ID validator and credit code validator run exactly these rules) while colliding with any real registered individual at astronomically small probability — because it is paired with no name, no address, no device, and points at no real database record.
The test data generator implements precisely this: batch ID numbers, mobile numbers, credit codes, bank cards (public test ranges), plates, and VINs — every item passes its validator. The engineering self-test loop: generate a batch → feed your form → deliberately corrupt one digit → confirm the validator catches it. That loop surfaces both failure classes — checks too loose and checks too strict — and the latter, in production, locks real users out.
The red lines (unlawful regardless of how the data was generated)
- Impersonating someone on a real service — passing platform KYC with a synthetic ID number is false-identity registration even if the format is perfect
- Defeating real-name verification — synthetic phone numbers against SMS-verified services fail at the gateway anyway, and attempting it is the violation
- Phishing / fraud material — forged ID-style pages with real-looking names
- Generating name-number pairs — the moment a synthetic number binds to a real name, the "test data" becomes a personal-information forgery tool
One self-check question: if this dataset left the test environment and touched a real system or a real person, what happens? "Nothing" means you are safe; anything else means stop.
How this pairs with the validators
Synthetic generation and format validation are two faces of one standards library: the generator guarantees "what it produces is well-formed," the validators guarantee "what is malformed gets caught." This format tool suite (IDs, credit codes, plates including new-energy positions, VINs, HK/TW/MO IDs) derives both directions from one rule set — data from the generator always passes the validators, and samples the validators reject are exactly what the generator never emits.
Test data is engineering hygiene; compliance is a legal floor. The two never conflict — the only conflict is with the habit of copying production straight into staging.
Related Tools
Related Articles
PIPL PIA Self-Check: The Five Triggering Scenarios, Three Assessment Elements, and a Working Framework
China's Personal Information Protection Law (Articles 55-56) requires a prior Personal Information Protection Impact Assessment (PIA) for five categories of processing. This post covers the triggers, the statutory three-element report structure, the three-year retention requirement, and a data-map → risk-matrix → controls framework. With the MLPS/miping/PIA three-pillar relationship.
MLPS 2.0 (China's Cybersecurity Multi-Level Protection Scheme) Self-Check: GB/T 22239, Layer by Layer
A practical self-check framework before China's MLPS level-protection evaluation (dengbao ceping): the filing workflow, high-frequency items across the five technical and five management layers, level-2 vs. level-3 differences, how MLPS relates to the cryptographic assessment (miping), and a remediation order ranked by points-per-effort.
China Generative AI Filing Checklist: Algorithm Registration vs. Large-Model Launch Filing
Offering generative AI services to the public in China means passing two filings: the algorithm registration under the Deep Synthesis Provisions and the large-model launch filing under the Interim Measures for Generative AI Services. This post maps the trigger conditions, the materials framework, the corpus and security-assessment pain points, and a self-check order. Practical guidance; always defer to the latest CAC templates.