Ox Alpha Model In 2026: The Stealth AI That Topped The Charts

Julian Goldie — founder, AI Profit Boardroom
By Julian Goldie · 8 min read
Get The AI Profit Stack Join AIPB →
🎯 1,000+ done-for-you AI agent workflows 📅 5 live coaching calls / week with me 🛡️ 7-day refund + 30-day ROI guarantee 👥 3,000+ AI operators inside

The ox alpha model was Z.ai's GLM-5.3-Flash running under a stealth alias — the company confirmed it on 26 August 2026, ending a week of speculation about the anonymous frontier model that appeared on OpenRouter with a million-token context window and a price of zero. So if you have been searching for who built Ox Alpha, whether it is still free, and what happens to your setup now the mask is off, here is the complete answer: what shipped, what it costs today, and what I found when I put it through our own testing.

📺 Watch: Z.ai Just Revealed the Truth About Ox Alpha

🔥 Get the Agent OS as a free bonus: AI Profit Boardroom members get the full Agent OS zip, prompt libraries, daily tutorials and weekly live coaching calls. → Get inside

What The Ox Alpha Model Actually Was

From 20 to 26 August 2026, an unbranded model listed as stealth/ox-alpha ran on OpenRouter. Per the coverage of the reveal, it offered a 1,048,576-token context window, up to 131,072 output tokens, and accepted text, images and video as input — completely free during the stealth window. Nobody knew who made it, which was exactly the point: Z.ai says it ran the model anonymously to gather real-world feedback before the official rollout, without brand bias skewing the results.

The community did not stay fooled for long. As I covered on my GLM 5.3 Flash free guide, testers fingerprinted the mystery model back to the GLM family within about 48 hours of it appearing. The behaviour, the tokeniser quirks, the response patterns — they all pointed the same way. When Z.ai's announcement landed on 26 August, it confirmed what the fingerprinting suggested: Ox Alpha was GLM-5.3-Flash, the first natively multimodal entry in the GLM-5 line.

The scale of the experiment is what makes this story remarkable. According to OpenCode's live usage data as reported around the reveal, the model had processed roughly 16 trillion tokens across 221,000 unique users and over 5 million sessions by 23 August — making it the number two model on that platform by recent usage while still wearing a fake name. That is not a soft launch. That is a full-scale public stress test conducted in broad daylight.

Ox Alpha Model Specs: What Z.ai Actually Shipped

Per the announcement, GLM-5.3-Flash — the ox alpha model's real identity — is a mixture-of-experts model with 320 billion total parameters and 18 billion active per token, built on a hybrid architecture combining sparse and linear attention. Z.ai says this design significantly reduces compute and memory-cache requirements, which is how a 320B model ends up priced like a small one. It is natively multimodal from the ground up: text, images and video in, text out.

SpecDetail (per the announcement)
ArchitectureMixture-of-experts, 320B total / 18B active parameters
Context window1,048,576 tokens
Max output131,072 tokens
Input typesText, image, video
LicenceMIT — open weights
Stealth window20–26 August 2026, free on OpenRouter

Two details matter most if you run agents rather than chat windows. First, the model supports tool calling and JSON-formatted outputs, which is the difference between a model you talk to and a model you build on. Second, the weights are released under the MIT licence — genuinely open, commercial-friendly, and yours to self-host if you have the hardware to serve a 320B-A18B model.

Want the exact agent stack I run new models through the day they drop? Inside the AI Profit Boardroom you get my full Agent OS zip, the model-testing workflows, daily tutorials and five live coaching calls a week — everything I used to test this model is in there. Steal my model-testing setup here.

Is Ox Alpha Still Free?

This is the question everyone types into Google, so let me answer it precisely. The stealth-window free ride ended when the reveal happened on 26 August. The stealth/ox-alpha listing was the temporary alias; the model now lives under its real name. But free-ish routes remain, and I rank them here best first.

  1. Number 1 — the open weights. The MIT licence means you can download the weights and run them yourself, permanently and commercially, at no licence cost. The catch is hardware: serving a 320B-A18B model locally is not a laptop job, as I explained when I tested running GLM 5.2 locally.
  2. Number 2 — the launch discount. Per the announcement, list pricing is 0.15 dollars per million input tokens and 0.50 dollars per million output — and a 50 percent launch promotion running until 9 September 2026 halves that to roughly 0.075 and 0.25. Not free, but close enough that a heavy month of agent usage costs less than a coffee.
  3. Number 3 — free tiers on routers. Aggregators periodically run free or promotional variants of new open-weight models. Availability changes weekly, so check the current listings — I keep my router notes updated in my Hermes agent OpenRouter guide.

For the full breakdown of every free route, including the self-hosting maths, my dedicated free-access guide goes deeper than this page — this one is about what the ox alpha model was and what the reveal means.

📺 Watch: This NEW Chinese AI Model Is Seriously Powerful

Ox Alpha Benchmarks And My Own Testing

On the official side, Z.ai says GLM-5.3-Flash beats GLM-5.2 across its internal evaluations while operating at roughly a tenth of the flagship price. The announcement also makes a claim that got the industry talking: the model runs entirely on Chinese AI chips. I will describe that neutrally — it is a supply-chain statement with real implications for pricing and availability, and it explains how Z.ai can sustain numbers this aggressive.

On my side, I do not publish anyone's benchmark table without running my own. Every model that enters my stack goes through Goldie Bench, our own testing suite built around the tasks that actually make money — landing pages from a single prompt, long-context document work, and agentic tool-calling runs. In our own testing, the pattern matched what the stealth-window crowd found: this model is startlingly fast for its quality tier, the million-token context genuinely holds up on retrieval tasks, and the tool-calling is reliable enough for unattended agent work. It sits in the same conversation as the other frontier Chinese releases I have covered in my Chinese AI models overview — and against the recent Tencent HY4 preview, Flash trades raw scale for speed and cost in a way that suits agent workloads better.

What The Reveal Means If You Were Using Ox Alpha

If you had agents pointed at the stealth listing, the practical migration is simple: the alias retires, the real model name takes over, and you move from a temporary free endpoint to a paid-but-cheap one. Your prompts and workflows carry over — it is the same model underneath. The bigger strategic point is about how releases now happen. Z.ai effectively ran a week-long public beta with a quarter of a million users before anyone knew the brand. Expect more stealth drops from other labs, because the feedback this generated was clearly worth more than a launch-day press cycle.

My advice is to build your stack so model identity is a config value, not an architecture decision. I run everything through my own Agent OS, where swapping the underlying model into a Hermes agent takes two clicks — which is exactly what let me move my test agents from the stealth endpoint to the official one in minutes, as I showed in my Kimi K3 Hermes agent setup when I did the same swap in the other direction.

How I Am Actually Using It

Right now the ox alpha model — under its real name — is my default cheap-and-long-context option for research agents: scraping and summarising competitor content, reading entire documentation sites in one pass, and drafting first-cut landing pages. The multimodal input means one agent can watch a screen recording and write the tutorial for it. At promo pricing, an overnight batch of a few hundred agent runs costs pennies, which changes the calculus on what is worth automating at all.

A quick answer to the other question people search: yes, there is a proper API now. The stealth listing was the anonymous preview; since the reveal, the model is available through Z.ai's own platform and through the usual aggregators under its real name, with the pricing above. If your first contact with the model was the free stealth endpoint, the API version behaves the same in my experience — same context handling, same tool-calling reliability — you are simply paying the (very small) bill now.

The reveal also confirmed something I keep repeating: the interesting story in AI right now is not the household names, it is the pace and pricing coming out of the open-weight labs. A model that spent a week as the second most used on a major platform under a fake name is the clearest proof of that you will ever get.

Ready to put models like this to work while you sleep? Join 3,000+ operators inside the AI Profit Boardroom — you get the full Agent OS as a free bonus, 1,000+ done-for-you agent workflows, daily tutorials and weekly live coaching calls with me. Get inside before the next stealth drop lands.

Real wins from inside the AI Profit Boardroom

See all 3,000+ members →
AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot

Ready To Join The #1 AI Community?

Join 3,600+ entrepreneurs inside the AI Profit Boardroom. Get 1,000+ plug-and-play AI agent workflows, daily coaching, and a community that holds you accountable.

Join The AI Community →

7-Day No-Questions Refund • Cancel Anytime

← Back to all posts