What Is Kaleb? 2.4T Stealth Model Explained

Elena Rossi

Elena Rossi

AI Adoption Analyst

Published: July 19, 2026
Kaleb stealth model reference card

TLDRKaleb is a stealth LMArena model, community-linked to Alibaba's 2.4T Qwen3.8-Max-Preview. Specs, access, and open questions.

What Is Kaleb? The 2.4T Stealth Model Community-Tied to Qwen3.8-Max-Preview

Kaleb is the codename of a stealth large language model that surfaced on the arena.ai / LMArena battle interface in mid-July 2026, alongside a second unknown entry called torenia-alpha. Community testers have linked Kaleb to Alibaba's newly announced Qwen3.8-Max-Preview, a 2.4-trillion-parameter model that the Qwen team says is "second only to Fable 5" among current frontier systems. The identification is based on output fingerprints, refusal patterns on China-political prompts, and timing, and is unconfirmed by Alibaba.

Key Takeaways

  • Kaleb first appeared as an anonymous entry on the LMArena battle interface on July 18, 2026, spotted by users including @Lentils80 and @Conor_D_Dart.
  • When asked, Kaleb claims to be Claude, but Chinese-language signals and political-topic handling point to a Chinese origin, per community testing.
  • On July 19, 2026, Alibaba's Qwen team publicly announced Qwen3.8-Max-Preview at 2.4 trillion parameters, positioning it "second only to Fable 5."
  • Community consensus, including a widely-shared read from @pankajkumar_dev, is that Kaleb is Qwen3.8-Max-Preview in stealth.
  • Leaked internal benchmarks circulated by @Lentils80 put the model within 12.6 points of Claude Fable 5 on coding and ahead of Claude Opus 4.8 Max by ~5 points on 400 agentic tasks — unverified and preview-stage.
  • Open weights for the full Qwen3.8 line are promised "soon" after the preview, with global rollout targeted for end-of-July 2026.

What Is Kaleb?

Kaleb is a stealth model slot on the LMArena battle interface, where anonymous models are pitted against each other in blind comparisons. It has no public model card, no vendor documentation, and no confirmed API endpoint under the name "Kaleb." Its existence is attested only by users who happened to draw it in arena matches and shared screenshots.

The name is a placeholder, similar to past stealth entries used by frontier labs to test models pre-launch. Kaleb's behavior — pretending to be Claude, refusing certain China-related prompts in ways characteristic of Chinese-trained models — has become the primary evidence for community attribution. As tester @Lentils80 put it, the model "pretends [to be] Claude but Qwen signals + Chinese political handling confirm Chinese origin."

The timing of Kaleb's appearance — one day before Alibaba's Qwen3.8-Max-Preview announcement — is the second load-bearing piece of evidence. No other Chinese lab announced a launch in that window.

Kaleb at a Glance

FieldValue
CodenameKaleb (stealth)
Suspected identityQwen3.8-Max-Preview (community-attributed, unconfirmed)
DeveloperAlibaba Qwen team (suspected)
TypeLarge language model, chat
Parameters2.4 trillion (per Qwen3.8 announcement, if attribution holds)
ModalityText; multimodal details not yet confirmed
Context windowNot yet confirmed
PricingNo standalone price; Qwen3.8-Max-Preview shown on Alibaba Token Plan at reported $6 / $20 / $70 tiers (Singapore)
AvailabilityLMArena battle mode (stealth); Qwen3.8-Max-Preview via Alibaba Token Plan, Qoder, QoderWork
Open weightsPromised for Qwen3.8, timing "soon" — not yet released
LicenseNot yet confirmed

How Kaleb Works and What Makes It Different

Kaleb's public footprint is limited to arena outputs, so architecture claims come from the associated Qwen3.8-Max-Preview announcement rather than from Kaleb itself. The Qwen team describes Qwen3.8 as a 2.4-trillion-parameter system and positions it just below Anthropic's Claude Fable 5 among current frontier models. If the Kaleb-equals-Qwen3.8 mapping holds, the model represents Alibaba's largest disclosed parameter count to date.

Two behaviors distinguish Kaleb in arena testing. First, users report unusually strong frontend and web-design output; @aniruddhadak's hands-on tests highlighted the quality of generated interface code. Second, the model's identity-cloaking — insisting it is Claude when asked directly — is more polished than typical stealth entries, suggesting deliberate system-prompt work rather than a raw base model.

torenia-alpha, the second stealth entry that arrived alongside Kaleb, remains unidentified. Community speculation ranges from a smaller Qwen sibling to an entirely separate lab's entry. A distinguishing detail some testers noted: torenia-alpha's knowledge cutoff appears to be late 2025.

What You Can Do With Kaleb

Practical use of Kaleb, today, means running it in LMArena battles and hoping you draw it. That is not a workflow. The more useful path is treating Kaleb as a preview of what Qwen3.8-Max-Preview will look like on Alibaba's own surfaces, where the same or a closely related model is served under its official name via the Token Plan, Qoder, and QoderWork.

Early testers have focused on three areas:

  • Frontend generation — HTML/CSS/JS scaffolding and component work, where community screenshots show polished results.
  • Agentic coding — leaked internal numbers cover a 400-task Cowork suite, suggesting Alibaba is targeting agent workloads.
  • General chat with strong reasoning — arena win-rate signals, while noisy, point to competitive quality against Opus-tier models.

For teams that want an open-weight Chinese frontier model available right now rather than waiting for the Qwen3.8 weights drop, our writeup of what Kimi K3 offers covers the current alternative.

How Kaleb Compares

The comparison numbers below come from leaked internal benchmarks attributed to Alibaba sources and shared by @Lentils80. They are preview-stage, not independently verified, and cover a 400-task agentic / Cowork suite rather than public evals.

ModelReported delta vs Qwen3.8-Max-Preview (agentic, 400 tasks)
Claude Fable 5+12.6 (Fable 5 ahead on coding; ~73.2% identical results)
Claude Opus 4.8 Max−5 (Qwen3.8-Max-Preview ahead)
Kimi K3−8 (Qwen3.8-Max-Preview ahead)
GLM-5.2−17.1 (Qwen3.8-Max-Preview ahead)
Qwen3.7-Max (prior generation)−44.4 (Qwen3.8-Max-Preview ahead)

Treat these as directional. For a grounded, public comparison of one of the neighbors on this table, see our analysis of Kimi K3 vs Claude Opus 4.8.

Availability: How to Access Kaleb

Kaleb itself has no dedicated endpoint. It surfaces randomly in LMArena's battle mode, where two anonymous models generate side-by-side responses. There is no way to select it by name.

The Qwen3.8-Max-Preview model that Kaleb is community-linked to is available through Alibaba's own channels: the Alibaba Token Plan, and the Qoder and QoderWork products. Pricing screenshots shared by testers show $6, $20, and $70 personal tiers on the Singapore Token Plan, with preview credits reportedly discounted and further nighttime cuts noted. These numbers are user-reported and may change once the model exits preview. Full open weights for the Qwen3.8 line are promised by Alibaba but do not yet have a public release date; the Qwen team described the timeline as "soon."

If you want to build against a comparable open-weight frontier chat model available today via API, Kimi K3 is the closest currently-shipping analogue.

What We Don't Know Yet

Because Kaleb is stealth and Qwen3.8-Max-Preview is fresh, most specification details remain open:

  • Official confirmation that Kaleb (and torenia-alpha) map to Qwen3.8-Max-Preview or its variants.
  • Context window size, exact architecture (dense vs MoE, active-parameter count), and multimodal scope beyond text.
  • Full public benchmark scores against Fable 5, Opus 4.8, Kimi K3, and GPT-5.6 — the leaked numbers are internal and preview-stage.
  • Firm release date and license terms for the promised Qwen3.8 open weights.
  • Whether Token Plan preview pricing survives to general availability.
  • What torenia-alpha actually is.

Frequently Asked Questions

What is Kaleb?

Kaleb is the codename of a stealth large language model spotted on the arena.ai / LMArena battle interface in mid-July 2026. Community testing links it to Alibaba's Qwen3.8-Max-Preview, though Alibaba has not officially confirmed the mapping.

Who made Kaleb?

Kaleb is widely believed to come from Alibaba's Qwen team, based on output fingerprints and its arrival window matching the Qwen3.8-Max-Preview launch. The model pretends to be Claude when asked, which is common stealth behavior, so the attribution rests on community signals rather than a vendor statement.

Is Kaleb open source?

Kaleb itself, as a stealth arena entry, is not downloadable. Alibaba has said Qwen3.8 will ship as open weights soon after the Max-Preview period, so if the community identification holds, weights are expected but not yet released.

How can I try Kaleb?

Kaleb appears in the LMArena battle mode as an anonymous option, so you can only encounter it by chance during blind comparisons. The associated Qwen3.8-Max-Preview is served through Alibaba's Token Plan, Qoder, and QoderWork surfaces.

How much does Kaleb cost?

Kaleb has no standalone pricing because it is a stealth arena entry. Community screenshots of the Token Plan tier for the associated Qwen3.8-Max-Preview show $6, $20, and $70 personal tiers in Singapore, with preview credits reportedly discounted.

What is torenia-alpha, and is it related to Kaleb?

torenia-alpha is a second stealth model that appeared on arena.ai around the same time as Kaleb. Testers speculate the two are related Qwen variants, but neither the relationship nor the identity has been confirmed.

How does Kaleb compare to Claude and Fable 5?

Leaked internal benchmarks attributed to Alibaba show the Qwen3.8-Max-Preview trailing Claude Fable 5 on coding by about 12.6 points while beating Claude Opus 4.8 Max by roughly 5. These numbers are unverified and drawn from a 400-task internal suite, not public evaluations.

What to Watch Next

Three signals will resolve most of the open questions here. First, an official Qwen post confirming (or denying) the Kaleb attribution — the Qwen team has been public about Qwen3.8-Max-Preview but silent on stealth. Second, LMArena publishing an Elo score for Kaleb once its testing window ends, which will replace the leaked internal numbers with a controlled measurement. Third, the actual Qwen3.8 open-weights release, which will let independent labs run the benchmarks that currently exist only as screenshots. This page will be updated in place as each lands.

Building similar chat and coding capabilities? On kie.ai you can try Kimi K3, Claude Fable 5, and Claude Opus 4.8.

Elena Rossi

About Elena Rossi

Elena watches developer chatter and early adoption signals to gauge which releases gain real traction.

View all posts by Elena Rossi