Ox Alpha Stealth Model: Signal vs Noise
Kenji Tanaka
Inference Systems Writer

TLDRA free 1M-context stealth model appeared anonymously. What fingerprinting confirms and what stays unverified.
Ox Alpha Stealth Model: Signal vs Noise on the Anonymous 1M-Context Release
A frontier-class model appeared with no company, no system card, and no press release. It arrived as a single line on a routing platform: stealth/ox-alpha, free, 1M context, released August 20, 2026. Within a day, coding agents pushed billions of tokens through it and the timeline turned into a detective game about who built it.
TLDR Ox Alpha is an anonymous "stealth" model that surfaced free on OpenRouter and OpenCode around August 20, 2026, advertising a 1,048,576-token context window and text, image, and video input. Community fingerprinting with an open-source tool points strongly at Zhipu's GLM family, but no lab has claimed it. The viral benchmark numbers are small self-reported tests, not audited results. Treat the specs as advertised, the identity as a strong inference, and the scores as unverified.
Key Takeaways
- Ox Alpha appeared as a free stealth model under the identifier
stealth/ox-alpha, with no lab attached as of August 22, 2026. - The advertised specs are a 1M-token context window, up to ~128K–131K output, and text/image/video input.
- Open-source tokenizer and error-code fingerprinting points at the GLM-5.X family; every other lab's best match is 2 of 4 tokenizer probes.
- A leaked system prompt confirms the anonymity is deliberate, instructing the model to name only "an undisclosed organization."
- The circulating benchmark scores come from small single-user tests, not audited leaderboards.
- The 100-trillion-tokens-per-day capacity figure is a relayed claim, and several developers openly doubt it.
What Was Actually Shipped
Two first-party listings anchor everything else. OpenRouter and OpenCode both put the model live around August 20, 2026, and the concrete details come from those cards rather than from the reaction thread.
The provider listing describes a reasoning model built for coding, sustained agentic work, and production workloads, with a context length of 1,048,576 tokens, as documented on the AIHubMix model page. A community relay of the OpenCode announcement adds the access terms: 1M context, multimodal input, zero data retention, and near-unlimited usage for roughly one week, per Chubby's summary.
Here is the hard-fact list, each drawn from a listing or a first-party relay:
- Model identifier:
stealth/ox-alpha, listed under the provider name "stealth." - Context window: 1,048,576 tokens, i.e. the full 1M ceiling.
- Maximum output: reported at roughly 128K–131K tokens, subject to route.
- Input modalities: text, image, and video, described on the provider card.
- Price: $0.00 per million tokens in and out during the preview.
- API shape: OpenAI-compatible Chat Completions, using a Bearer key.
One usage signal is worth pinning. On the OpenCode data page, ox-alpha ranked #3 across the prior week's usage with about 2.0% of observed volume and roughly 7.1T tokens moved, drawing about 134K unique users. That is not a benchmark. It is adoption, and it explains why the identity question caught fire so fast.
Data Retention Claim is the one spec that carries real risk for builders: the anonymous provider states it does not train on prompts or completions, which is not the same as promising nobody can read them. Zero-data-retention is a stated policy, not an audited guarantee.
The Fingerprinting Evidence: How the Community Traced It
The most useful reporting here is not a hot take. It is infrastructure probing. When Ox Alpha appeared, one developer collected the community's scattered tricks into a single open-source tool called modelprint, which runs nine infrastructure probes against any OpenAI-compatible endpoint and compares fingerprints side by side.
Tokenizer Fingerprinting is the core method: because tokenizers are built per lab, counting how a model splits pinned strings of CJK text, emoji, and code reveals the family behind an anonymous API. The tool's day-one run put the mystery model against twelve suspects. Only the GLM family matched all four normalized tokenizer counts. Every other lab's best result was 2 of 4.
The reported breakdown was blunt: 6/9 total probes and 4/4 tokenizer matches for z-ai/glm-5.3, 5/9 for z-ai/glm-4.7-flash, and 2/9 or lower for GPT, Qwen, Kimi, DeepSeek, MiniMax, Gemini, Grok, and Claude. That is the single strongest piece of evidence in the whole discourse.
A second independent probe chain, relayed in an OpenCode community thread, reported the same conclusion through Error-Code Family matching: sending reasoning_effort: "none" triggered a verbatim [1210] rejection identical to Z.ai's GLM-5.3 API, and forcing an image error returned a Chinese-language parse message, suggesting a Z.ai server rather than a US proxy. On the limited evidence so far, Ox Alpha appears to sit inside the GLM-5.X family — but no lab has confirmed it, and fingerprints identify a serving stack, not a product name.
The anonymity itself is deliberate, not accidental. A leaked system prompt shared by Pliny the Liberator instructs the model to identify strictly as "ox-alpha, developed by an undisclosed organization," and to refuse any other identity. The same author later attributed it to "Zai, GLM-5.X family."

Source: @elder_plinius
Why This Matters for Builders
The stealth-model pattern is now a recognized release channel, not a one-off. A third-party writeup counted Ox Alpha as the fifth anonymous release in about six months, with the previous four all resolving to Chinese labs. That cadence changes how engineers should read these events.
First, the economics are the story. A model advertising 1M context at $0.00 during a free window is buying real-world evaluation data at scale. The 7.1T tokens moved in a week is the payment: it is a distributed load test and a preference-collection run, dressed as a giveaway.
Second, the identity investigation has become reproducible engineering. Two years ago, "which lab is this" was a vibe check. Now it is a nine-probe script anyone can run against an endpoint, which matters far beyond stealth models — the same tool caught a mainstream API silently swapping the model behind an old alias. That is a durable capability for any team that depends on a provider honoring what it advertises.
Third, the safety surface is unresolved. One tester reported no cybersecurity guardrails during a hands-on run, and the zero-retention claim is unaudited. For any workload touching secrets or proprietary code, a free anonymous endpoint is a poor default regardless of code quality.
Ox Alpha vs GLM 5.3: What the Signal Says
Because the bundle names GLM 5.3 as the leading identity theory, the sharpest comparison is between the advertised Ox Alpha listing and what the community reports about GLM 5.3. This is not a measured head-to-head. It is a mapping of one advertised spec sheet against community claims, with hedges intact.
- Context window: Ox Alpha advertises 1,048,576 tokens. No public GLM 5.3 context figure appears in this signal set, so the comparison is unverified on the GLM side.
- Modalities: Ox Alpha lists text, image, and video. Multiple observers note that public GLM 5.3 is reported as text-only, which is why the theory is specifically an "unreleased multimodal GLM variant," per Lumina's read.
- Tokenizer fingerprint: 4/4 normalized token counts match
z-ai/glm-5.3in the modelprint run. This is the strongest link and the one hardest to fake, because a serving template cancels out under normalization. - Error-code family: Ox Alpha reportedly returns Z.ai-style
[1210]and[1301]codes, matching what GLM-5.3's API is documented to emit in community testing. - Pricing: Ox Alpha is $0.00 during the preview. No standard GLM 5.3 price appears in this signal set, so post-preview cost is unknown.
Community impressions run positive but split. Teortaxes wrote that it reads like "another 5.x + vision," tighter and more professional, and that he could see it "trading blows" with a top model. Others were harsher: Bindu Reddy noted it "spins a lot" and called it "yet another chinese model." Neither is a benchmark.
How to Evaluate It Yourself
The circulating scores deserve skepticism. A self-reported 30-task comparison put Ox Alpha at 10/10 on coding, 15/19 on adversarial tests, and 3/10 on FrontierMath, while a different model led overall. A separate account claimed 96.6% on Terminal Bench 2.1 across 89 tasks using best-of-five, with no methodology attached. Neither is an audited leaderboard.
If you want a number you can trust, generate it. A practical protocol:
- Pin a fingerprint first. Run the tokenizer and error-code probes yourself before you trust any identity claim in a thread. If a thread tells you "Ox Alpha is GLM," that person is theorizing until the probes agree.
- Bound the task. Send relevant files and constraints, ask for a concrete output format, and validate with your normal tests. A 1M context window does not replace review.
- Watch the failure modes. One tester reported a repository-wide review capped at four parallel subagents despite requesting ten, with repeated tool failures. "Spins a lot" showed up in more than one report; measure wall-clock time, not just output quality.
For a stable, documented baseline to run the same eval against, a comparable long-context chat model such as Kimi K3 gives you a fixed reference point while Ox Alpha's identity and pricing stay in flux.
What We Know vs. What We Don't
What we know:
- Ox Alpha appeared as a free stealth model on OpenRouter and OpenCode around August 20, 2026, under the identifier
stealth/ox-alpha. - It advertises a 1,048,576-token context window and up to roughly 128K–131K output.
- It is described as accepting text, image, and video input.
- It is currently listed at $0.00 per million tokens for input, output, and cache reads.
- A leaked system prompt confirms the model is instructed to identify only as "an undisclosed organization," making the anonymity deliberate.
- Open-source fingerprinting matched all four normalized tokenizer probes to the GLM family, with error codes echoing Z.ai's API.
What we don't:
- No lab has officially claimed Ox Alpha; the GLM-5.X attribution is a strong inference, not a confirmation.
- No official benchmarks have been published; the viral scores are small single-user tests.
- The 100-trillion-tokens-per-day capacity is a relayed claim, doubted by several developers, not a demonstrated throughput.
- Parameter count, training data, and exact architecture are undisclosed.
- The pricing and terms after the free week are unresolved, with no standard price stated.
- Whether the zero-data-retention policy holds in practice is unaudited.
What to Watch Next
Three signals will resolve most of the open questions. Watch for an official model card or a lab stepping forward to claim the release, which every prior stealth model in this cadence eventually did. Run your own fingerprint and coding eval before trusting the identity claim or the 96.6% Terminal Bench number. And check whether the $0.00 listing survives the announced week or converts to standard pricing, since the post-preview terms are the one detail no source has pinned down.
Building similar long-context coding or agentic chat workflows? On kie.ai you can try Kimi K3, Gemini 3.8 Flash, and Claude Opus 4.8.
About Kenji Tanaka
Kenji follows latency, throughput, and pricing signals to separate hype from shipped capability.
View all posts by Kenji Tanaka