Skip to content
Main Site News Console

Anonymous Niulai LLM Revealed to Have Zhipu Lineage; Some Suspect It Was Trained on Open-Source GLM

· 量子位
国内AI

Tokenizer, Video Encoding, and API Errors: The Full Ox Alpha Investigation

Hengyu, reporting from Aofeisi Temple
QbitAI

Whose model is the “Niulai” model that has taken the world by storm?

Over the past two days, the AI community has practically turned it inside out.

At present, the two most widely discussed candidates for the model’s developer are Zhipu and Google.

Those who suspect Zhipu have already started examining its tokenizer, measuring video token usage, and hunting for API errors—almost as if they were conducting a digital DNA test on the model.

Those who suspect Google are mainly focused on recent posts by DeepMind employees, as well as the Gemini 3.5 Pro that has yet to make an official appearance.

The “Niulai” model is actually called Ox Alpha.

Since “ox” means “牛” in Chinese, and the “Niulai” meme has recently taken the Chinese internet by storm, Chinese netizens gave it the affectionate nickname “the Niulai model.”

On August 20, it suddenly appeared on OpenRouter as an anonymous model. OpenRouter only said that it came from a third-party provider that was temporarily keeping its identity confidential.

It has a 1.0486-million-token context window and can generate up to 131,000 tokens in a single response. It can process text, images, and video, and is specifically designed for code, long-running Agents, and complex reasoning.

Then OpenCode announced that Ox Alpha was available through OpenCode Go, offering almost unlimited free use for several consecutive days, without counting against users’ existing Go quotas (public discussions even mentioned service capacity on the order of 100 trillion tokens per day).

Developers then descended on it with codebases, terminals, and all kinds of tricky tasks.

And began asking:

Come on, Niulai—whose model are you, really?

Why These Two Companies?

People suspect that the Niulai model comes from either Zhipu or Google because its deep understanding of Chinese, long-context processing, and multimodal integration all show characteristics associated with the models developed by these two companies.

It has also revealed similar technical fingerprints in its underlying engineering details.

Why Zhipu?

Zhipu is the leading favorite in this guessing game.

Netizens have traced clues from the Niulai model’s tokenizer and video encoding method all the way to its API service layer.

Let’s start with the tokenizer.

A tokenizer splits input into tokens that a model can process. Different model families often use different vocabularies and tokenization logic, so carefully designed text can leave behind relatively stable “fingerprints” in the token count.

Some testers compared Ox Alpha with several other models using different prompts and found that Ox Alpha’s token counts showed a highly consistent relationship with GLM-5.3.

In one widely circulated test, the same passage containing a mixture of Chinese and English, code, and emoji was counted as 68 tokens by both GLM-5.3 and GLM-5.2, while Ox Alpha reported 143.

The additional 75 tokens could be explained almost exactly by a hidden System Prompt appended externally to the model input.

Another tester used 25 different prompts for verification and observed a similar fixed offset.

The multimodal fingerprint is even more interesting.

A text model can imitate another model’s response patterns through distillation, post-training, or a System Prompt. But video encoding and token budgets involve the engineering design used before multimodal inputs enter the model.

One researcher created four controlled test videos and fed them separately to Ox Alpha and several multimodal models, observing how many input tokens the systems allocated to the videos.

The results showed that Ox Alpha and GLM-5V-Turbo had highly similar token budgets across multiple characteristics.

These included the following: changes in video frame rate had little effect on the sampling strategy; as video duration increased, token usage grew at approximately 147 tokens per second; and the way token usage scaled with per-frame resolution was also quite similar.

Across the four controlled videos, the two models’ additional token budgets even matched on an item-by-item basis.

The MiMo, Qwen, and GLM-4.6V models included in the comparison displayed different characteristics.

There were also traces left at the API service layer.

Someone in the community deliberately sent invalid parameters to Ox Alpha and captured an error message resembling [1210] The temperature parameter is illegal.

Similar error codes, parameter restrictions, and responses mixing Chinese and English were also considered highly similar to those found in Zhipu-related services.

Some users even claimed to have spotted residual text such as “trained by Z.ai” in the model’s reasoning output.

However, this claim currently comes primarily from community screenshots and secondhand accounts. The evidence is clearly weaker than the tokenizer and video-token tests above, so for now it is better treated as an Easter egg.

There is another piece of context that makes the “it’s Zhipu” theory seem even more plausible.

In February this year, OpenRouter launched another anonymous model called Pony Alpha, which prompted a similar round of speculation.

In the end, OpenRouter revealed the answer: Pony Alpha was an early test version of GLM-5.

So when Ox Alpha appeared, many people had the same immediate reaction:

Zhipu, is it you again?

Google Suddenly Becomes a Favorite in the “Identify the Ox” Contest

Zhipu’s clues had been leading the race, but over the past two days, the plot suddenly turned toward Google.

Indeed, Ox Alpha’s public product profile makes it easy to associate the model with Gemini.

A million-token context window, native image and video input, support for long-running Agent tasks, and an emphasis on complex code engineering and tool calling—all of these have been major areas of investment for Google across recent generations of Gemini.

And Google happens to have another Gemini Pro model that has yet to be officially announced.

During Google I/O this year, Google suggested that Gemini 3.5 Pro was “Coming next month,” but its release date was repeatedly delayed. To this day, the industry is still waiting for a new Gemini Pro flagship model. There have also been multiple rounds of leaks concerning its release schedule, internal testing, and product adjustments.

In particular, recent social media activity from Google and DeepMind employees has given netizens plenty of room for speculation.

For example, Google DeepMind researcher Evan Otero posted “Gemini” in a related discussion, followed by: “What if Ox Alpha is the friend we’ve encountered all along?”

Of course, neither post explicitly stated that “Ox Alpha is Gemini.”

But against the backdrop of the internet-wide frenzy over the model’s identity, the posts were quickly interpreted by the community as vague posting.

In the social media discussion compiled by Techmeme, a number of developers also noticed that members of the Gemini team had suddenly begun talking much more frequently about Gemini and next-generation models. Some consequently connected the Niulai model with Gemini 3.5 Pro—or even Gemini 4.

The internet quickly split into two camps:

One side waved tokenizer analyses and video-token bills, shouting, “This is obviously GLM.”

The other pointed to DeepMind employees’ posts and said, “Google, stop pretending.”

At present, there is no public evidence supporting this theory. But it does show that identifying a model’s developer based solely on its external behavior is becoming increasingly difficult.

Especially as distillation, model merging, post-training, and third-party inference services become more widespread, “who trained the model” and “where the model is deployed” may even be two separate questions.

One Free Ox Has Drawn Programmers From Around the World

Putting the identity mystery aside, the Niulai model is interesting enough on its own.

First, its specifications.

A 1M-token context window theoretically allows a single Prompt to contain an entire large code repository, hundreds of pages of documentation, or a lengthy record of Agent runs.

Its 131K maximum output is also far beyond what is typical for ordinary chat scenarios.

Combined with image and video input, tool calling, and a mandatory reasoning mode, Ox Alpha was clearly designed from the outset for code Agents and long-running tasks.

According to publicly available data from OpenCode, the Niulai model has already accumulated usage on the order of 250B tokens, with approximately 8,500 unique users and more than 110,000 completed sessions.

After trying it himself, Stripe CEO Patrick Collison also commented: “It’s very impressive.”

A few small-scale coding tests pushed it even further into the spotlight.

In a community test involving 10 real-world software engineering tasks, Ox Alpha completed eight, for a score of 80%.

On the same set of tasks, Fable 5 scored 65%, GLM-5.3 scored 62%, and GPT-5.6 Sol scored 52%.

It is important to stress that this was a small-sample community test involving only 10 questions. It was not an official Benchmark and is far from sufficient to prove that Ox Alpha’s overall capabilities surpass those of these flagship models.

Meanwhile, plenty of people have pushed back against the hype.

Abacus.AI CEO Bindu Reddy said publicly that after testing Ox Alpha, they found its performance below expectations, with some metrics reaching only roughly the level of an earlier-generation model.

Wharton professor Ethan Mollick also said it was “not all that impressive.”

Other developers tested it on their own private Benchmarks and concluded that its performance was average at low reasoning intensity. In visual tasks, some also found it clearly weaker than the Gemini series.

So for now, the most accurate description of Ox Alpha may be:

An anonymous model that performs exceptionally well on certain coding and Agent tasks, boasts extremely aggressive specifications, and has produced highly divergent evaluation results.

Its free access, combined with its mysterious identity, has given it an extra boost of attention.

One More Thing

Another joke making the rounds is that it could be Cursor retraining a Zhipu model in-house.

And then putting an ox head on it (?)

As for the answer, it probably will not remain hidden for long.

After all, judging from the storylines surrounding OpenRouter’s recent Alpha models, once the anonymous period ends, someone will eventually have to come out and lead the ox home.

Mooo—

References:

[1]
https://openrouter.ai/api/v1/models

[2]
https://opencode.ai/docs/zen/

[3]
https://www.reddit.com/r/singularity/comments/1vufbx1/i_fingerprinted_ox_alpha_same_tokenizer_as_glm53/

[4]
https://www.businessinsider.com/ox-alpha-ai-model-mystery-2026-8

[5]
https://openrouter.ai/models?fmt=cards&input_modalities=video&q=amazon

[6]
https://ox-alpha-evidence-production.up.railway.app/