China Generative AI Filing Checklist: Algorithm Registration vs. Large-Model Launch Filing
Two filings — don't conflate them
Serving the public with generative AI in mainland China typically involves two gates:
| Algorithm registration | Large-model launch filing | |
|---|---|---|
| Basis | Provisions on Algorithmic Recommendations / Deep Synthesis Provisions | Interim Measures for Generative AI Services |
| Object | Providers of deep-synthesis / recommendation services | The generative AI service itself |
| Timing | Before launch / within 10 working days after (deep synthesis) | Before opening to the public |
| Platform | National algorithm filing system | Provincial cyberspace administration |
If you only wrap a purchased API but serve the public under your own brand, the filing obligation sits on your head regardless of whose model runs underneath.
Materials framework (practical; defer to official templates)
1. Entity and basics
- Business license, legal-representative materials
- Algorithm/product owner and contacts
- Service form description (web / app / mini-program / API)
2. Algorithm security self-assessment report (the core)
- Algorithm principles and model-scale description
- Corpus compliance: source legality, intellectual property, personal-information basis (ties into the PIA framework)
- Generated-content safety: interception mechanisms, keyword libraries, human review workflow
- Refusal and correction mechanisms (complaints, rumor channel)
- Data and model security: training-data cleaning records, model artifact protection
- Incident response plan (handling flow and SLA for unlawful content)
3. Security assessment and testing
- Third-party security assessment (required in several provinces)
- Stress and bias test records
- Generated-content sampling review records
4. Commitments and disclosure
- Truthfulness commitment letter (signed by the legal representative)
- AI-content labeling — prominent identification of generated text/images/audio is a hard technical requirement under the Deep Synthesis Provisions
Pain points, by frequency
- Corpus provenance unexplained — crawled data without authorization chains. Fix: a corpus-source ledger + commercial/open-source licensing evidence
- Interception fails live testing — stale keyword libraries, no jailbreak-test records. Run your own red-team round before the defense
- No AI-content label — output-side product changes are required, not optional
- Report and reality diverge — the report claims human review; no staffing schedule exists. Inspectors ask for execution records (same "evidence over intent" rule as the MLPS checklist)
- Training data contains personal information without basis — cross-check with PIPL; supplement separate consent or dedup/anonymization evidence
Suggested self-check order
- Pin down the service form (own model / fine-tuned open-source / API wrapper) — it decides the responsibility boundary and materials emphasis
- Corpus ledger and live content-safety testing first (slowest to fix)
- Report writing (follow the official template section by section; detail beats brevity)
- Pre-communication with the provincial CAC office (most provinces accept it and it saves a round trip)
Implementer's note
One lesson from writing this site's content transfers directly to filing materials: reviewers do not want declarations ("we have a mechanism") — they want verifiable specifics: keyword-library size, sampling ratio, review SLA in hours. Quantify what can be quantized, evidence what can be evidenced. And one cross-point: if your AI service processes user prompts containing personal information, that flow belongs in the PIA data map — the filing and the PIPL assessment feed each other.
Related reading
Related Tools
Related Articles
Post-Quantum Migration Self-Check: From Algorithm Inventory to Hybrid TLS — 12 Things to Do in 2026
A practical self-check framework for PQC migration: why three overlapping regulatory timelines make 2026 the starting line, how to build a cryptographic bill of materials (CBOM), ranking by HNDL exposure, enabling X25519+ML-KEM hybrid TLS, and what crypto-agility means in practice. With a 12-item checklist.
PIPL PIA Self-Check: The Five Triggering Scenarios, Three Assessment Elements, and a Working Framework
China's Personal Information Protection Law (Articles 55-56) requires a prior Personal Information Protection Impact Assessment (PIA) for five categories of processing. This post covers the triggers, the statutory three-element report structure, the three-year retention requirement, and a data-map → risk-matrix → controls framework. With the MLPS/miping/PIA three-pillar relationship.
Passkeys, Explained: Why the Server Can No Longer Hold a Copy of Your Password
How passkeys (WebAuthn/FIDO2) work: public-key challenge-response replacing shared secrets, why phishing resistance is structural (origin binding), what servers storing only public keys actually means, synced vs. device-bound passkeys, and the three enterprise adoption hurdles (enrollment, multi-device, recovery). With a capability table against passwords and TOTP.