Skip to content
Crypto2026-10-04Begin3 min read

China Generative AI Filing Checklist: Algorithm Registration vs. Large-Model Launch Filing

Two filings — don't conflate them

Serving the public with generative AI in mainland China typically involves two gates:

Algorithm registrationLarge-model launch filing
BasisProvisions on Algorithmic Recommendations / Deep Synthesis ProvisionsInterim Measures for Generative AI Services
ObjectProviders of deep-synthesis / recommendation servicesThe generative AI service itself
TimingBefore launch / within 10 working days after (deep synthesis)Before opening to the public
PlatformNational algorithm filing systemProvincial cyberspace administration

If you only wrap a purchased API but serve the public under your own brand, the filing obligation sits on your head regardless of whose model runs underneath.

Materials framework (practical; defer to official templates)

1. Entity and basics

  • Business license, legal-representative materials
  • Algorithm/product owner and contacts
  • Service form description (web / app / mini-program / API)

2. Algorithm security self-assessment report (the core)

  • Algorithm principles and model-scale description
  • Corpus compliance: source legality, intellectual property, personal-information basis (ties into the PIA framework)
  • Generated-content safety: interception mechanisms, keyword libraries, human review workflow
  • Refusal and correction mechanisms (complaints, rumor channel)
  • Data and model security: training-data cleaning records, model artifact protection
  • Incident response plan (handling flow and SLA for unlawful content)

3. Security assessment and testing

  • Third-party security assessment (required in several provinces)
  • Stress and bias test records
  • Generated-content sampling review records

4. Commitments and disclosure

  • Truthfulness commitment letter (signed by the legal representative)
  • AI-content labeling — prominent identification of generated text/images/audio is a hard technical requirement under the Deep Synthesis Provisions

Pain points, by frequency

  1. Corpus provenance unexplained — crawled data without authorization chains. Fix: a corpus-source ledger + commercial/open-source licensing evidence
  2. Interception fails live testing — stale keyword libraries, no jailbreak-test records. Run your own red-team round before the defense
  3. No AI-content label — output-side product changes are required, not optional
  4. Report and reality diverge — the report claims human review; no staffing schedule exists. Inspectors ask for execution records (same "evidence over intent" rule as the MLPS checklist)
  5. Training data contains personal information without basis — cross-check with PIPL; supplement separate consent or dedup/anonymization evidence

Suggested self-check order

  1. Pin down the service form (own model / fine-tuned open-source / API wrapper) — it decides the responsibility boundary and materials emphasis
  2. Corpus ledger and live content-safety testing first (slowest to fix)
  3. Report writing (follow the official template section by section; detail beats brevity)
  4. Pre-communication with the provincial CAC office (most provinces accept it and it saves a round trip)

Implementer's note

One lesson from writing this site's content transfers directly to filing materials: reviewers do not want declarations ("we have a mechanism") — they want verifiable specifics: keyword-library size, sampling ratio, review SLA in hours. Quantify what can be quantized, evidence what can be evidenced. And one cross-point: if your AI service processes user prompts containing personal information, that flow belongs in the PIA data map — the filing and the PIPL assessment feed each other.