A production cluster with no Nvidia in it

Z.ai published a technical account on 17 September of how it built a production inference service for GLM-5.3-Flash from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for the model now runs on that system, the company said. It did not name the chipmakers.

That omission matters, because the claim underneath is a capability claim about China’s domestic silicon: that a frontier-scale model can be served commercially without US accelerators, at a cost the company says is competitive.

The model being served

GLM-5.3-Flash was released on 26 August. It has 320 billion total parameters with 18 billion active, a one-million-token context window, and is the first model in the GLM-5 series to take image and video input as well as text.

Before release it was tested anonymously under the name ox-alpha. Z.ai says it became the most-used offering on OpenCode and OpenRouter within a week, processing more than 62 trillion tokens in six days.

A technician holding an etched circuit panel at a workbench
The company says an agent running on its own model carried out much of the optimisation work. Photograph for illustration. Willquezada · pexels · Pexels License

An agent doing the infrastructure work

The part of the account that is not about chips is about who did the work. Z.ai says much of the optimisation was carried out by an internal Infra Agent running on GLM-5.3 — handling analysis, hypotheses and code changes — while human engineers set the objectives, the system boundaries and the risk assessment.

To make that workable the team built what it calls dense feedback: correctness tests, runtime logs, execution traces, runtime events, microbenchmarks and end-to-end metrics, assembled so a proposed change could be checked locally instead of requiring a full deployment after every edit.

The numbers, and where they come from

Z.ai reports going from initial model adaptation to production readiness in under two weeks, with end-to-end throughput ultimately three times the starting baseline, and says hardware utilisation and per-token cost reached levels comparable to mainstream Nvidia GPUs.

Every one of those figures is the company’s own, about its own system, and none has been independently verified. The account also lists what made it hard: limited on-chip memory capacity and bandwidth, a new model architecture, the million-token context, multimodal requests, and an immature software ecosystem in which kernel support was incomplete and engineers had to infer behaviour that should have been documented.

Network cables and a patch panel in a data centre
The cluster is described as holding more than 100,000 accelerators. Photograph for illustration. Brett Sayles · pexels · Pexels License

The export-control context

US export controls have kept the most capable Nvidia parts out of China for three years, and Chinese labs have been building around what they can buy at home. Huawei this week pulled its next Ascend accelerator forward by three quarters, to the first quarter of 2027, and Chinese chipmakers have been raising prices as high-bandwidth memory runs short.

Analysts quoted in coverage of the report have suggested the cluster is likely a mix of Huawei Ascend parts and hardware from other vendors, but that is inference from outside, not something Z.ai has confirmed.

What to watch

Two things would turn this from a company claim into a checkable fact: Z.ai naming the accelerators, and an independent measurement of serving cost per token against Nvidia-based providers. Until then, the load figures are the firmest part of the story — a model this heavily used has to be running somewhere.