Reuters reported on October 9, 2026 that research firm SemiAnalysis analyzed 857 model releases from nine leading Chinese artificial intelligence developers between 2021 and September 15, 2026, and found that only 31 of them (3.6%) publicly disclosed model-specific safety evaluation results. Only 9 of those (1.1%) were published at or before release; the remaining 813 releases carried no safety disclosure at all.

How the study was conducted

The study covered Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot, Z.AI, MiniMax, and StepFun — four large tech giants and five startups. Each release was cross-checked against the developer's model card, release notes, and technical reports.

The study period spans more than five years — from 2021 to September 15, 2026; during that time the nine companies released models 857 times in total.

The counting methodology was strict: the researchers counted only publicly announced, specific results tied to a clearly named model. These include findings on harmful output, jailbreak resistance, toxicity, privacy, refusal behavior, and dangerous capabilities. Generic claims such as "trained for safety" or "evaluated" were not counted — only results with specific findings or figures were accepted.

To be counted, a result had to be tied to that exact model and include a specific finding or figure — a criterion that automatically excluded generic claims like "trained for safety" or "evaluated." The researchers checked each release individually against the developer's model card, release notes, and technical reports.

The numbers behind the 3.6%

Of the 857 releases, only 31 contained a model-specific safety result — 3.6% of all releases. In only 9 of them was the result published at or before the release, meaning the share of releases with a safety evaluation at launch does not exceed 1.1%. The remaining 813 releases — 94.9% of the total — contain no safety results disclosed by the developer. Put differently, only about one in 28 releases publicly disclosed a model-specific safety result, and models whose results came out at or before release number about one in a hundred.

The full breakdown: of the 31 results, 9 were published at or before release, 16 were documented after release (median delay: 42 days), and for 6 the timing or model attribution could not be determined; another 10 releases carried evaluation claims without figures, 3 had only press or investor information, and the remaining 813 had no disclosure at all.

So even of the 31 disclosed results, 16 were documented late, after release (median delay: 42 days), and for 6 the timing or model match was never established — meaning that in most cases where disclosure happened, users could not see a model's safety evaluation when it launched.

Another finding of the report: no major Chinese developer has released a frontier text model with openly published results of dangerous-capability tests covering cyber, biological, and loss-of-control risks. In other words, the capabilities considered most dangerous in top models have never been announced as having undergone open testing.

"China's real approach to AI safety is based on speed, not safety," SemiAnalysis researchers concluded in their report.

China's rules: apps controlled, model frontier open

According to SemiAnalysis, China's AI safety governance framework recognizes leading risks but imposes no mandatory requirements tied to model capabilities. That is, risks are noted in official documents, but they do not automatically impose obligations on developers depending on model size or capability.

Beijing's mandatory rules instead mainly regulate apps and their impact on users — what is released and how it affects people is controlled, but there is no requirement for how the model itself must be tested or whether results are disclosed. The report's main conclusion follows from this: the Chinese regime is strict at the app level but operates without mandatory guardrails at the leading-model level.

In practice, this means that strict control at the app level cannot compensate for the absence of mandatory guardrails at the level of the most advanced models — the report's authors describe the Chinese regime precisely this way: strict at the app level, but without mandatory guardrails at the leading-model level.

Comparison with US labs

The report provides no comparable figures for US developers — SemiAnalysis counted only Chinese companies' releases and did not conduct the same full count for US labs. Direct numerical comparison of the two sides is therefore impossible. However, as Reuters notes, OpenAI, Anthropic, and Google DeepMind have published safety reports, system cards, or model cards when releasing some of their major frontier models. That is, some Western labs have a practice of openly publishing safety documentation in their most important releases, while Chinese developers lack such a practice systematically.

This difference shows not in numbers but in practice: some Western labs have a practice of openly publishing safety documentation in their most important releases, while Chinese developers lack such a practice systematically.

SemiAnalysis published the original study on October 8, 2026 under the title "Beijing Will Not Pace the Frontier: China's Speed-First AI Safety Regime." The report's main thesis is reflected in its title: the authors describe China's real approach to AI safety as speed-first. Reuters covered the study's findings the following day, October 9.