Skip to content
Main Site News Console

Zhipu GLM-5.3-Flash Launches with Homegrown Compute from SenseTime’s Big Device

· 量子位
国内AI

Domestic Heterogeneous Computing Helps Bring Cutting-Edge Intelligence to the Masses

On the evening of August 26, Zhipu officially launched and open-sourced GLM-5.3-Flash (320B-A18B). This is the first natively multimodal model in the GLM-5 series, with 320B total parameters and capabilities that surpass GLM-5.2. On the globally authoritative Artificial Analysis Intelligence Index (AA Intelligence Index), GLM-5.3-Flash scored 57, placing it among the world’s frontier models and matching the score of Anthropic’s most popular model, Claude Opus 4.8. GLM-5.3-Flash features an architecture designed for ultra-low cost. Combined with Zhipu’s latest 30T-token multimodal pretraining corpus, it delivers stronger performance with fewer computing resources.

Powered by advanced domestic heterogeneous hybrid inference technology, SenseTime’s SenseCore provides GLM-5.3-Flash with large-scale domestic computing support and token services, establishing another landmark example of the large-scale commercial deployment of domestic computing power.

Behind this collaboration is SenseTime’s core capabilities, built around “one model, one Token factory, and one agent management system.” The Token Factory plays a crucial role in the large-scale production of intelligent capabilities: through the joint optimization of multimodal models and SenseCore infrastructure, it coordinates different chips, clusters, energy resources, and delivery nodes to continuously produce higher-quality tokens at lower cost. At the same time, the Token Factory and foundation models form a bidirectional feedback loop, enabling better adaptation to the iterative needs of natively multimodal models. SenseCore’s efficient support for the launch of Zhipu’s GLM-5.3-Flash is a powerful testament to this closed-loop capability.

For a long time, the market has viewed domestic computing power as lacking cost-effectiveness and difficult to deploy commercially at scale. Before its official release, however, GLM-5.3-Flash underwent large-scale testing on OpenCode and OpenRouter under the anonymous model name Ox-Alpha, known in the Chinese developer community as “Niu Lai.” Token usage reached as high as 62T, with all of this request traffic powered by domestic chips. Compared with the initial baseline under the same hardware conditions, GLM-5.3-Flash, running on a domestic chip cluster, achieved a threefold improvement in end-to-end service performance. Its hardware efficiency and per-token cost reached levels comparable to mainstream NVIDIA GPUs.

This practice demonstrates that domestic computing power is now capable of efficiently and economically supporting the inference demands of cutting-edge foundation models in large-scale, real-world business scenarios.

As the AI infrastructure that best understands foundation models, SenseCore fully embraces domestic technology and continues to drive domestic computing power from “usable at individual points” toward “large-scale commercial deployment.” During this year’s WAIC, SenseCore addressed the pain points of large-scale commercial deployment of domestic computing power by systematically tackling the issue across three dimensions: cost-effectiveness, adaptability, and energy efficiency.

In terms of cost-effectiveness, SenseCore moves beyond the traditional focus on single-chip performance comparisons. Through advanced heterogeneous hybrid inference technology, it assigns domestic chips with different architectures and roles according to their distinct computing and bandwidth requirements at different inference stages, enabling each to play to its strengths and work together efficiently. Overall inference cost-effectiveness reaches 1.25 times that of NVIDIA’s H-series GPUs. Compared with homogeneous inference using domestic chips, token throughput increases by approximately 2.5 times at the same cost.

In terms of adaptability, SenseCore significantly lowers the barriers to applying domestic computing power through low-level operator optimization, multi-chip parallelism tuning, and toolchain adaptation. In terms of energy efficiency, SenseCore has introduced a compute-energy coordination Agent and redefined TPW (Tokens Per Watt), making it a new value metric for AIDC.

In addition, by integrating diverse computing resources and continuously optimizing inference engines and chip performance, the SenseCore Token Factory is building a large-scale, highly efficient, and low-cost token supply capability. In July, its average daily token service volume reached 2.42 trillion. By the end of 2026, this figure is expected to exceed 10 trillion tokens per day, with annual token volume growing 25-fold.

Looking ahead, SenseCore will continue investing in token services powered by domestic computing resources, building AI infrastructure that delivers high performance, low cost, and scalable deployment, and driving domestic computing power from isolated technological breakthroughs toward the flourishing of an integrated industrial ecosystem.

This article was provided by SenseTime and reprinted by QbitAI with authorization. All views expressed are those of the original author.