Skip to content
Main Site News Console

The Mysterious “Niulai” Model Is Indeed Zhipu! GLM’s First Native Multimodal Model, Powered by Homegrown GPUs

· 量子位
国内AI

< img id=“wx_img” src=“https://www.qbitai.com/wp-content/uploads/imgs/qbitai-logo-1.png” width=“400” height=“400”>

2026-08-27
00:48:12
Source: [QbitAI](https://www.qbitai.com/)
GLM-5.3 Flash
>

Jin Lei, reporting from Aofeisi

QbitAI | Official WeChat account: QbitAI

At last, the mysterious model “Ox Alpha” that has been making waves online over the past few days has been revealed!

And, as expected, it comes from a Chinese player—

It is GLM-5.3 Flash, just released and open-sourced by Zhipu—and it is also the first natively multimodal model in the GLM 5 series.

During the few days that “Ox Alpha” was trending, quite a few people had already been testing the Ox Alpha version in practice.

One of the more popular experiments was building an interactive SpaceX Raptor rocket engine 3D webpage from scratch.

For example, after using the same prompt to have “Ox Alpha” complete this task twice over the course of a few days, blogger Tim Jayas shared this reaction:

It feels like “Ox Alpha” is a model that can keep learning on its own.

As it happens, we also obtained access to the GLM-5.3 Flash beta and built the same project. After we entered the prompt, no further operation was needed:

After waiting a while, the 3D Raptor engine project was ready. Take a look at the result:

Video:

https://mp.weixin.qq.com/s/wNQAWeUacfB8a1F_kQZe_g

Judging from the comparison, well, there is little more to say: the 3D Raptor engine produced by GLM-5.3 Flash is clearly more refined.

So why has the “Ox Alpha” model become so popular?

We can get a glimpse of the answer from posts by various official sources.

On OpenRouter, for example, “Ox Alpha” climbed to the top of the rankings on its very first day and set a new record for daily token usage:

OpenCode said that Ox Alpha ended DeepSeek’s 56-day run at the top of the OpenCode rankings immediately after its launch:

After Zhipu officially claimed “Ox Alpha,” we also discovered several other highlights of GLM-5.3 Flash.

In terms of model size, GLM-5.3 Flash has just 320B parameters, yet its capabilities comprehensively surpass those of the much larger 753B GLM-5.2.

What’s more, according to the latest AA leaderboard, GLM-5.3 Flash has already scored 57 points, placing it among the global frontier models and tying Claude Opus 4.8.

In terms of pricing, GLM-5.3 Flash is priced at one-tenth of the regular GLM-5.3 rate. Its limited-time discount brings it down to one-twentieth of GLM-5.3 and one-fortieth of Opus 4.8, making it cheaper than DS-V4-Flash.

There is also one more thing worth celebrating—

The 62T tokens mentioned above were all processed on domestically produced chips! The mysterious model that overseas developers had been chasing for a month turned out to be a thoroughbred Chinese “ox.”

So, as usual, let’s dive into an in-depth hands-on test~

Zhipu GLM Finally Has Eyes

As we mentioned earlier, Zhipu’s GLM-5.3 Flash is the first natively multimodal model in the 5 series.

So for our first round of testing, we focused on multimodality.

For example, we gave GLM-5.3 Flash a video featuring multiple speakers and asked it to determine who was speaking based on the original video, then add precisely synchronized subtitles.

The prompt was:

First identify the different speakers, then create subtitles that combine speaker-specific colors with a karaoke-style read-along effect: keep each speaker’s background color fixed, while highlighting the word or phrase currently being read. Use a high-contrast palette of cyan-blue, warm yellow, coral red, and white, with the current word transitioning smoothly from low brightness to its highlighted color. Distinguish between the main speaker, host, and brief interjections; do not list background music, applause, or ambient voices as separate speakers. The subtitles must be realigned based on the final timeline, and no edited clip may reuse the original video’s previous timestamps. Review the entire video before making decisions to avoid inferring the later structure solely from the opening sample. Keep cuts natural and prioritize clear dialogue. Place subtitles within the title-safe area and use outlines or background panels to ensure readability against complex backgrounds.

This task is a significant test of a model’s multimodal processing capabilities. It first needs to understand the video, correctly match the people, and, most importantly, generate accurate subtitles.

Now let’s see how GLM-5.3 Flash performed:

Video:

https://mp.weixin.qq.com/s/wNQAWeUacfB8a1F_kQZe_g

One noteworthy detail is that GLM-5.3 Flash was able to spot a name tag that appeared only once in the footage, and use it to establish the relationship between the voice, the name, and the person.

This shows that GLM-5.3 Flash delivered a solid multimodal performance on this task.

If adding subtitles seems a little too easy, then let’s “feed” GLM-5.3 Flash an entire classic film, The Myth, and ask it to directly produce a movie commentary video.

The prompt was:

Please edit this complete film into a Chinese movie commentary video. Organize the main plot and character relationships, highlight key conflicts and emotional turning points, ensure that the commentary remains consistent with the original film, and generate voice-over narration and Chinese subtitles.

Here is the final video:

Video:

https://mp.weixin.qq.com/s/wNQAWeUacfB8a1F_kQZe_g

Compared with the previous subtitling task, GLM-5.3 Flash had to “watch” a much longer video this time, maintain an understanding of the entire plot, and select key clips that were both coherent and important.

Judging from the final result, GLM-5.3 Flash was able to understand and execute all the crucial steps involved in creating a movie commentary video on its own. The final presentation was also more than satisfactory.

Next, in addition to multimodality, let’s also put its coding capabilities to the test.

The basic task was to give GLM-5.3 Flash a complete product design mockup and ask it to turn the image into an interactive product interface.

Here is the design mockup (scroll vertically to view the full image):

The prompt was:

Implement a complete mobile shopping app based on this UI design mockup. Reproduce the original design’s color scheme, fonts, images, cards, and page layouts as closely as possible, and turn the main pages and shopping flow shown in the image into a fully interactive experience. Implement the following features: Home/Discover, product listings, filters, product details, reviews, favorites, shopping cart, delivery, addresses, payment, order completion, store map, personal center, and settings.

Here is the final interactive result:

Video:

https://mp.weixin.qq.com/s/wNQAWeUacfB8a1F_kQZe_g

Once again, a whole series of features could be interacted with smoothly as requested in the prompt.

For an even more complex task, Zhipu officially demonstrated the results of a 3D Blender scene-building task that GLM-5.3 Flash ran independently for 12 hours:

Video:

https://mp.weixin.qq.com/s/wNQAWeUacfB8a1F_kQZe_g

So, the current situation with GLM-5.3 Flash is this: as long as you give it enough time, it can work through complex tasks on its own, at a low cost.

The Architecture Is Entirely New, Too

After seeing the results of the hands-on tests, you may be wondering:

How can a 320B model achieve results like these?

The answer is actually hidden behind the word “Flash.”

This time, Zhipu began optimizing efficiency at the model-architecture level itself.

GLM-5.3 Flash has 320B total parameters, but only 18B active parameters. Its layer count has also been reduced from the 92 layers used during the GLM-4.5 era to 45 layers. Yet despite this slimming-down, its capabilities exceed those of the larger previous-generation GLM-5.2. Combined with the latest 30T-token multimodal pretraining corpus, Zhipu’s goal was to compute less without doing less.

The most important change was made to the attention mechanism. GLM-5.3 Flash uses a hybrid architecture combining linear attention and sparse attention.

Linear attention handles local information, while sparse attention uses a lightweight indexer to retrieve the truly relevant global context. With an ultra-long 1M-token context, the model does not need to perform intensive all-to-all computation among every token.

Zhipu also added an IndexPool, compressing the indexer’s original four cached vectors into one. As a result, compared with GLM-5.3, GLM-5.3 Flash reduces attention computation by 3.01× and cuts the KV Cache by 4.44×.

This also explains a clear impression from the earlier tests: although it bears the name “Flash,” it does not feel like a “lightweight version” when handling long-running tasks.

This is especially true of the “eyes” it has finally grown. Zhipu developed a dedicated data-synthesis pipeline for Visual Coding, allowing the model to inspect the final pages, interactions, and 3D scenes while carrying out a task, then continue making modifications based on visual feedback.

That is why, in tasks such as reproducing design mockups and 3D modeling, it can write, inspect, and revise all at the same time.

In addition, the real-world online requests currently handled by GLM-5.3 Flash are all processed on domestically produced chips behind the scenes.

To support multimodality and 1M-token contexts on domestic accelerator chips, Zhipu split multimodal encoding, prompt prefill, and token-by-token decoding into an Encode-Prefill-Decode decoupled architecture that can be scheduled, scaled, and expanded independently. It also implemented a range of low-level optimizations, including Layer Split and mixed cache quantization.

Ultimately, compared with the initial baseline on the same hardware, end-to-end serving performance improved threefold, while the cost per token reached a level comparable to mainstream NVIDIA GPUs.

Made in China, Open-Sourced in China—Because We’re Confident

Seen in this light, the sudden overseas popularity of “Ox Alpha” over the past few days takes on a different meaning.

Before it officially claimed the model, neither OpenRouter nor OpenCode’s overseas developers knew who Ox-Alpha belonged to. They simply found it useful, kept feeding it prompt after prompt, and propelled it to the top of the rankings.

Then, when Zhipu revealed its identity, everyone discovered there was another part to the story:

The model was made in China, and even the computing power supporting this wave of global real-world traffic was provided by Chinese hardware.

As mentioned above, during its anonymous testing period, Ox-Alpha set new usage records on both OpenRouter and OpenCode—and every one of those requests was powered by domestically produced chips.

The significance of this goes beyond “a Chinese model winning once again.” Frontier models have faced an increasingly practical problem: as their capabilities improve, they also become more expensive to use.

Especially now that agents are genuinely getting work done, a single task can run for dozens of minutes or several hours, with tokens flowing out as if a faucet had been left open. No matter how intelligent a model is, when the price remains high, users still instinctively start doing the math.

The answer GLM-5.3 Flash offers this time is simple: continue squeezing efficiency out of the architecture, reasoning, and computing infrastructure, and bring down the price of frontier capabilities along with it.

It scored 57 on AA, tying Claude Opus 4.8, yet costs only one-fortieth as much. Frontier models are finally starting to feel a little more like everyday consumables.

Finally, there is open source.

GLM-5.3 Flash has now been officially open-sourced worldwide and integrated with ZCode and the Open API. The model can run on domestically produced chips, while its weights are openly available. Developers can access them, and enterprises have the opportunity to deploy the model themselves.

The model, domestic computing power, and open-source ecosystem have truly come together this time.

So at this point, the name “Ox Alpha” suddenly seems rather fitting.

It gets the job done and eats little. And this time, the ox is ours, and so is the grass.

BigModel Open Platform: https://docs.bigmodel.cn/cn/guide/models/vlm/glm-5.3-flash

Z.ai: https://docs.z.ai/guides/vlm/glm-5.3-flash

GLM Coding Plan: https://bigmodel.cn/glm-coding

Blog: https://z.ai/blog/glm-5.3-flash

Open-source links: https://huggingface.co/zai-org/GLM-5.3-Flash https://huggingface.co/zai-org/GLM-5.3-Flash-BF16

*Copyright. Reproduction or use in any form without authorization is prohibited. Violators will be held legally liable.*