What Do We Know About Ox Alpha? Not Much.

An anonymous model drops on OpenRouter with frontier-level specs and no label.

I have always enjoyed a good mystery. The intrigue, the suspects, putting together pieces of a puzzle. But mystery models are a different story. They show up with impressive specs, no label, and a crowd of developers immediately trying to figure out who made them. It is part detective story, part early-access window, and part genuine risk-management question. I find them cautiously interesting.

Ox Alpha is the latest one. It dropped on OpenRouter on August 20, 2026 and by the next morning everyone in the developer community had an opinion. This model had the developer’s beehive buzzing. I decided to spend some time with it. Here is what I found, what I think, and where the honest answer is still I do not know.


What Just Dropped

On August 20, 2026 at 20:04 UTC, a model called Ox Alpha appeared on OpenRouter under the label stealth/ox-alpha. Anonymous third-party provider. No lab name. No paper. No press release. Just a model in the directory. I recently added OpenRouter to one of my projects and I have been impressed with the results so far.

It popped up around the same time on OpenCode. Still considering the use case for that tool. Reasoning is mandatory and always on. OpenRouter’s own description, pulled directly from the model page: “a reasoning model designed for coding, sustained agentic work, and production workloads.” Suited for “long-horizon software engineering, complex reasoning, and workflows that combine text with visual context. Sounds good right?”

That is everything officially stated on day one. Then the information pool starts to get a litle lower. Everything after that is either community testing or speculation. This post is going to be clear about which is which.


The Specs That Are Actually Confirmed

These come directly from the openrouter.ai model page, verified August 25, 2026.

From openrouter.ai/stealth/ox-alpha verified
Context window
1,048,576
tokens (~1M)
Max output
131,072
tokens
Input modality
Text, image, video
text output only
Throughput
25 t/s
tokens per second
Latency
4.51s
time to first token
Price
$0 / $0
during preview
Tool calling
Supported
JSON output supported
Parameters (est.)
~744B total
~40B active, MoE estimated

These estimates are from independent technical analysis, not from the provider. These numbers are amazing! The Mixture of Experts (MoE) which not a group of really smart people. Its an AI based technique that splits a large model into smaller specialized sub networks called “experts” that process data more efficiently.That’s a standalone blog topic on its own. The MoE structure would explain how the claimed 100 trillion tokens per day capacity is economically feasible at zero cost. Seems logical right? Confirmed fact? No. math. Let’s be 100% clear on that point, because that’s where we are.

One data point that caught my attennion before you build anything long: 25 tokens per second is on the slower end for sustained agentic work. If your workflow involves long multi-step loops, budget for the pace. It is not a dealbreaker, but it is not snappy either.


The Usage Numbers Tell a Story

Whatever questions remain about who built this, the adoption data is not ambiguous. These numbers are sourced directly from OpenRouter and OpenCode’s own public disclosures.

tokens in 3 days
11.6T
OpenRouter’s largest launch ever, 2.6x previous record
tokens via OpenCode
26T
across 4 days of free preview
unique users
327K
on OpenCode in 4 days
completed sessions
8.3M
average 25 sessions per user
Chart 1 of 3
Launch scale: Ox Alpha vs previous OpenRouter record
previous largest launch Ox Alpha (first 3 days)

Source: OpenRouter public disclosure. Ox Alpha processed 2.6x more tokens in its first 3 days than any prior model launch on the platform.

One important caveat from OpenCode’s own dashboard: 93% of input tokens were served from cache. The headline totals combine fresh inputs with context reused across long coding sessions. The number is real. What it represents is sessions, not raw new prompts.

What is also telling is where the traffic came from. The OpenRouter model page shows which apps sent the most tokens to Ox Alpha in its first days. This was not chat traffic. It was coding agents doing real work.

Chart 2 of 3
Who sent the most traffic to Ox Alpha on OpenRouter
token volume (trillions)

Source: openrouter.ai/stealth/ox-alpha, verified August 25, 2026. Every top sender is a coding agent, not a chat interface.

Teams running real coding agents routed real work to this model in its first 72 hours. That is a different signal than a viral social media moment. It does not tell you who built it or whether it holds up at production scale. But it does tell you developers treated it like a workhorse, not a toy.


What the Benchmarks Actually Show

The 80% DeepSWE number has been everywhere. Here is what that number actually is, and what came after it.

An independent developer ran Ox Alpha through a 10-task subset of DeepSWE shortly after launch and got roughly 80%. The tester flagged the sample size immediately: at 10 tasks, one task equals 10 percentage points. A near-miss in either direction moves the headline number significantly. That is not a criticism of the tester. It is just math.

/Additional sample data set showed that when it ran the full 113-task DeepSWE benchmark, the score came back between 58% and 63%. That puts Ox Alpha roughly in line with Claude Opus 4.8 on the same benchmark. Still a strong result for a free anonymous model. Not the frontier-beating headline the 10-task number suggested.

Chart 3 of 3
DeepSWE benchmark: the headline vs. the full picture
Ox Alpha (10-task community run) Ox Alpha (full 113-task run) other models (official leaderboard)

Ox Alpha has no official DeepSWE leaderboard entry. The 10-task and 113-task figures are community runs. Official scores sourced from deepswe.datacurve.ai (August 2026).

Ox Alpha has no official entry on the DeepSWE leaderboard as of August 25, 2026. The data points we have are community runs under conditions that do not match the official harness. Strong signal. Not a verdict.

The dev community is definitely on board with this model “word on the street” from the first few days of community testing is telling: Ox Alpha handles long multi-step instructions well, uses tools competently, and stays coherent deep into large contexts. That matches what the app traffic data suggests. It is doing the kind of work coding agents actually do, not just generating good-looking code snippets.


The Part Nobody Knows, and I Mean Nobody

The model’s creator is undisclosed. That is not a fine-print disclosure. Its the whole story.

The leading theory is Zhipu AI’s unreleased GLM-5.x series. On Manifold Markets, 63% of participants are betting on Z.ai as of August 24. Community fingerprinting found a constant 75-token difference from Zhipu AI’s GLM-5.3, a video-encoder match to GLM-5V-Turbo token-by-token, and an API error format consistent with Zhipu’s services. A deliberately malformed request to OpenCode’s Ox Alpha route returned a Java stack trace with an internal class name: com.wd.paas.api.domain.v4.chat.ChatCompletionRequest. That signature is consistent with Zhipu AI’s known infrastructure. Not confirmation. Evidence.

One complication: GLM-5.3 listed on OpenRouter two days before Ox Alpha appeared is text-only. Ox Alpha takes images and video. If the attribution holds, this is not GLM-5.3 itself but an unreleased multimodal sibling. The community idea box is bursting with names like GLM-5.3V or GLM-5.5. Those are speculation. Nothing has shipped under those names.

Identity theories as of August 24, 2026 (Manifold Markets)
Z.ai / Zhipu AI (GLM-5.x)
63%
Other / Unknown
12%
Xiaomi (MiMo)
~9%
Alibaba (Qwen) / Moonshot (Kimi)
~8%
OpenAI / Anthropic (fringe)
4% each

Ox Alpha is the fifth anonymous stealth release on OpenRouter in six months. The previous four all followed the same arc: anonymous debut, burst of free traffic, then the lab stepped forward. The playbook is established. Ox Alpha will almost certainly get named. We just do not know when.


The Security Questions I Cannot Ignore

Earlier reporting flagged a contradiction between OpenCode and OpenRouter on retention. Here is where things stand as of August 25, verified against both platforms directly.

OpenRouter’s model page now states clearly: prompts and completions are retained by the provider and are not used for training.

OpenCode’s route claims zero-day retention through its Go service. No storage, no training.

Two routes, two policies. They are not interchangeable. But the retention question is actually the smaller of the two issues here. The bigger ones are these.

Data and IP exposure. Ox Alpha’s operator is anonymous and retains every prompt on the OpenRouter route. If you feed it a real codebase, you are sending Supabase schema and queries, auth logic, service configs, and potentially API keys or environment variables if they end up in context, to an unknown party with no accountable identity, no published data-handling terms, and no way to request deletion. That is not a theoretical risk. That is the actual situation right now.

Supply chain and trust risk. Because you do not know who trained this model or what it was optimized for, you cannot rule out subtly biased or unsafe code suggestions making it into a production repo. Weakened auth checks. Insecure defaults. The kind of thing that does not look wrong at a glance but creates exposure downstream. With a known vendor you have a paper trail, a published safety posture, and some basis for audit. With an anonymous model you have none of that. The code still looks like code.

Your code prompt + context OpenRouter stealth/ox-alpha OpenCode Go / Zen endpoint Ox Alpha anonymous provider Prompts retained by anonymous provider no deletion path Zero retention claimed by OpenCode Go service only route A route B

Two routes, two retention policies. The model is the same. What happens to your prompt depends entirely on which endpoint you use.

This is not standard “be careful with sensitive data” soft touch. An anonymous operator retaining your prompts with no deletion path and no published terms is a specific, concrete risk. Say that out loud and sit with it. Its important. So is the possibility of subtly unsafe code from a model whose optimization targets you cannot verify. I call this the “free with cost factor”. Know exactly which route your requests are taking and what the terms of that route actually say before you run anything real through this.

Should You Try It?

My standard process before writing about any model is to add it to LLMCode Lab, run it through my testing stack, and build a real comparison profile against the models I already track. Its how I test, in additional to testing models live in my own projects.

I could not do that here. When a model’s identity is unconfirmed, there is nothing to anchor a proper profile to. I do not know who trained it, what data it saw, or what I am actually comparing. That is not a reason to skip writing about it. It is a reason to be honest that my read is narrow intentionally.

So: yes, try it. The free preview window is exactly the right moment to experiment. The confirmed specs are strong, the usage data shows real coding agents treated it seriously, and you are not paying to find out. Run it on isolated tasks. Test it against models you already know on OpenRouter.

What I would not do: route a real codebase through this. The retention risk and the supply chain risk are not hypothetical. An unknown operator holding your schema and auth logic with no deletion path is a real exposure. Code suggestions you cannot fully audit from a model whose optimization targets are unknown is a real exposure. Test it on problems that do not matter if they leak. Keep your production repos out of it until the identity is confirmed and the terms are clear.

What Would Actually Change My Mind

Provider identity disclosed. An official DeepSWE leaderboard entry run under the same harness as every other model. Data policy consistent across both routes in writing. And a confirmed identity that lets me run this through LLMCode Lab properly and give you a real comparison instead of a community signal read.

None of those have happened yet. Until they do, Ox Alpha sits in the same bucket as every other promising-but-unverified model I have seen come through OpenRouter: worth watching, worth experimenting with, not worth anchoring anything production on. I will spending more time with OpenRouter for sure.


Final Thoughts – For Now

A mystery model with frontier specs, 26 trillion tokens of usage in four days, and zero disclosure is exactly the kind of thing that gets developers excited. And exactly the kind of thing that deserves a slow second look before you route anything important through it. This one the reasons why I enjoy building and writing in this space, literally something new appears on the IT scene daily. What a great opportunity!

The specs are verified. The usage is real. The benchmark picture is more complicated than the headline. The identity is unknown. And if the pattern holds, we will know who built this within a week or two. When that happens, it will likely be a candidate for the LLMCode Lab so that’s great site to bookmark. I will update all the model details there.

Subscribe so you do not miss it. This space moves fast and I am going to keep showing up with the honest read, even when the honest read is: I do not know enough yet.

Sources: openrouter.ai/stealth/ox-alpha (verified August 25, 2026). OpenCode usage dashboard via opencode.ai/data/unknown/ox-alpha. OpenRouter public disclosure. Manifold Markets prediction market (August 24, 2026). Day.dev community benchmarks. DeepSWE leaderboard at deepswe.datacurve.ai.

Last verified: August 25, 2026.
The free preview window, data policy terms, provider identity, and benchmark results are all moving fast. This is all subject to change.Probably right after I hit submit.


Discover more from MsTechDiva

Subscribe to get the latest posts sent to your email.

Discover more from MsTechDiva

Subscribe now to keep reading and get access to the full archive.

Continue reading