What Is ox-alpha? Free 1M-Context Stealth Model
Priya Nair
AI Infrastructure Analyst

TLDRox-alpha: a stealth model on OpenRouter/OpenCode with 1M context, video input, and $0 preview pricing. Vendor unconfirmed.
What Is ox-alpha? The Free 1M-Context Stealth Model With Video Input
ox-alpha is an anonymous "stealth" AI model that appeared on OpenRouter and OpenCode on August 20, 2026, offering a 1M-token context window, text/image/video input, and free access during a roughly one-week preview. It is positioned as a reasoning model for coding, sustained agentic work, and production workloads. No company has publicly claimed it as of August 22, 2026, and community fingerprinting most often points to a Chinese lab, with Z.ai's GLM family the leading theory. It is served under the model ID stealth/ox-alpha.
Key Takeaways
- ox-alpha launched anonymously on OpenRouter and OpenCode on August 20, 2026, routed under the provider name "stealth."
- It carries a 1,048,576-token (1M) context window and accepts text, image, and video input, with function calling supported.
- Access is free during the preview: $0 input, $0 output, $0 cache reads, with OpenCode citing capacity of 100T tokens per day.
- The vendor is unconfirmed. Tokenizer probes and a Z.AI API error code seen by testers push most theories toward GLM-5.3; a minority argue for Gemini.
- Early hands-on reports praise code quality and multimodal reasoning but flag that it "spins a lot" and is slow to generate.
- On OpenCode's own usage tracker, ox-alpha ranked #3 last week with 7.1T tokens and ~2.0% of observed volume.
What Is ox-alpha?
ox-alpha is a stealth model: a frontier-class system released anonymously through API platforms so its unnamed lab can gather real-world feedback before an official launch. OpenRouter's listing describes it as a reasoning model designed for coding, sustained agentic work, and production use. The provider field reads "stealth," and OpenRouter states it only routes the requests.
The model is in an active preview state, not a formal release. There is no model card, no license, and no named vendor. What is confirmed comes from the platform listings and hands-on community testing; the origin story is still open. According to the OpenCode announcement, the free window runs about one week, and independent trackers already show heavy adoption during that span.
Its arrival fits a pattern. Coverage from a third-party explainer describes ox-alpha as the fifth anonymous stealth release in roughly six months, with the prior four all traced back to Chinese labs after the fact.
ox-alpha at a Glance
| Attribute | Detail |
|---|---|
| Developer | Not yet confirmed (community theories favor a Chinese lab / Z.ai GLM family) |
| Type | Reasoning model for coding, agentic work, production workloads |
| Modality | Text, image, and video input |
| Context window | 1,048,576 tokens (1M; ~1.05M on some trackers) |
| Max output | 128K tokens (per a third-party reference listing) |
| Model ID | stealth/ox-alpha |
| API format | OpenAI-compatible Chat Completions |
| Pricing | $0 during stealth preview; long-term pricing not yet confirmed |
| Availability | OpenRouter, OpenCode, and several agent harnesses (free, ~1 week) |
| License | Not yet confirmed (no weights released) |
| Knowledge cutoff | ~December 2025 (community-reported, unconfirmed) |
How ox-alpha Works / What Makes It Different
The defining traits are its 1M-Token Context Window and Multimodal Input across text, image, and video. Together these let it hold a full codebase, task history, and visual references in a single session, which is the core of its pitch for long-horizon work.
A second differentiator is behavioral: ox-alpha is a heavy reasoner. Multiple testers describe it as slow and prone to over-deliberation. "This model thinks way too much, but look at what it produced," wrote one user who generated a SpaceX Starship simulation inside Claude Code. Bindu Reddy's hands-on note lists the same tradeoff bluntly: "free, 1M context, multi-modal" as pros, and "spins a lot, may not be that useful in the real world" as cons, per her OpenRouter test.
The Vision-Grounded Understanding angle is what sets it apart from text-only coders. Teortaxes reported that ox-alpha could "quickly fix a lot of bullshit left behind" in a multi-week rendering-pipeline project, "leveraging precise vision-grounded understanding of effects of its interventions," per this post. That combination of long context plus visual grounding is the recurring reason testers cite for reaching for it on messy, real repositories.
What You Can Do With ox-alpha
Early demonstrations cluster around code-heavy and multimodal generation:
-
One-shot 3D and procedural scenes. One agent run produced a full three.js "dreamcore" world from a single prompt at 64,745 output tokens with no external assets, according to a community demo.
-
Voxel and creative rendering. Ivan Fioravanti generated a "Voxel Pagoda Garden Cherry Blossom Scene" with day/night transitions from one prompt, describing the model's attention to detail.
-
Multimodal game logic. Running ox-alpha through the Hermes agent on a Frogger prompt, the model added its own gameplay elements — "Multimodality in action!" — in this test.
-
Long-horizon repair. Fixing accumulated errors across a multi-week project with visual feedback loops, as described above.
These are individual demonstrations, not audited benchmarks. Treat them as capability signals rather than guarantees. If you want a documented long-context coding baseline to compare against, our analysis of DeepSeek V4 Flash covers a model with published prices and a similar 1M-context claim.
How ox-alpha Compares
Benchmark claims for ox-alpha are unaudited and contradictory, so treat this table as directional. One third-party 10-task DeepSWE comparison reported ox-alpha averaging 80%, ahead of several named rivals; separately, Bindu Reddy's evaluation placed it near two-generation-old models and called it "quite bad." Both are single-source claims.
| Model | Reported context | Pricing (per 1M) | Notes |
|---|---|---|---|
| ox-alpha | 1M–1.05M | $0 (preview) | Multimodal; unconfirmed vendor |
| DeepSeek V4 Pro | 1M | ~$0.66 in / $1.98 out | Published, audited scores |
| DeepSeek V4 Flash | 1M | ~$0.22 in / $0.66 out | Fast, low-cost |
The consistent, verifiable fact across sources is adoption, not quality: OpenCode's usage tracker ranked ox-alpha #3 over the prior week at 7.1T tokens, ~2.0% of observed volume, behind DeepSeek V4 Flash and Xiaomi's MiMo.
Availability: How to Access ox-alpha
ox-alpha is live during a free stealth preview through OpenRouter and OpenCode, plus several agent harnesses. The confirmed access facts:
- Model ID:
stealth/ox-alpha, sent in an OpenAI-compatible Chat Completions request. - Pricing: $0 input, $0 output, $0 cache reads during the preview.
- Window: roughly one week from the August 20, 2026 debut, per OpenCode.
- Capacity: OpenCode cited "near unlimited" usage and 100T tokens/day.
- Data handling: OpenCode stated zero data retention during the free period; this is a vendor claim, not independently verified. Multiple testers still advise keeping passwords, personal data, and proprietary code out during the window.
Because the model is anonymous and routed through third parties, availability, routing, and limits can change without notice. Confirm the live listing before you build a workload around it. If you need a hosted long-context multimodal chat model with stable, documented access today, you can call Gemini 3.5 Flash through the kie.ai API as a comparable option while ox-alpha's status settles.
What We Don't Know Yet
The open questions outnumber the confirmed facts:
- Who built it. The vendor is unconfirmed. Community fingerprinting via tokenizer patterns and a reported Z.AI API error code (
invalid zstd request body) points many testers to a GLM-5.3 variant, but this is inference, not disclosure. A few posts speculated Gemini; at least one directly disputed that ("Ox Alpha is NOT a new Gemini"). - Real benchmark standing. Reported scores range from "80% on DeepSWE" to "worse than last-generation models." None are reproducible, audited artifacts.
- Parameter count and architecture. Not published.
- Long-term pricing and stability. Undefined once the free window closes.
- Whether multimodal quality holds beyond individual demos.
Frequently Asked Questions
What is ox-alpha?
ox-alpha is an anonymous stealth AI model that appeared on OpenRouter and OpenCode on August 20, 2026, built for coding, agentic work, and production use. It offers a 1M-token context window, text/image/video input, and free access during a roughly one-week preview.
Who makes ox-alpha?
The company behind ox-alpha is unconfirmed as of August 22, 2026. OpenRouter routes it under the provider name "stealth" and no lab has publicly claimed it; community fingerprinting most often points to a Chinese lab, with Z.ai's GLM family the leading theory.
How much does ox-alpha cost?
ox-alpha is free during its stealth preview, with $0 input, $0 output, and $0 cache-read pricing on OpenRouter. Long-term pricing has not been published, and the free window is stated to last about one week.
Is ox-alpha open source?
ox-alpha is not open source as of August 22, 2026. No weights, license, or model card have been released, and the model is only accessible through hosted APIs during its stealth preview.
What is ox-alpha's context window?
ox-alpha has a 1M-token context window (listed as 1,048,576 tokens, or roughly 1.05M on some trackers). It also accepts text, image, and video input and supports function calling.
Is ox-alpha a Gemini or GLM model?
ox-alpha's identity is unconfirmed. Community fingerprinting based on tokenizer behavior and API error codes most often points to Z.ai's GLM family (a possible GLM-5.3 variant), while a minority of posts speculated about Google Gemini; none of these attributions are official.
How do I access ox-alpha?
ox-alpha is accessible through OpenRouter under the model ID stealth/ox-alpha and through OpenCode and several agent harnesses during the free preview. Availability, routing, and limits can change, so check the live listing before building on it.
What to Watch Next
Three signals will resolve most of the open questions. First, a vendor claim: the four prior stealth models were all unmasked within days, so watch for a Chinese lab (Z.ai or otherwise) stepping forward. Second, the free window's close — that is when real pricing, rate limits, and whether the 1M context and video input survive at scale become clear. Third, an audited benchmark to replace the contradictory single-account scores now circulating. This page will be updated as those land.
Building similar long-context multimodal models? On kie.ai you can try Gemini 3.8 Flash, Gemini 3.5 Flash, and GPT-5.6.
About Priya Nair
Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.
View all posts by Priya Nair