What Is Hy3? Tencent's 295B Open-Source MoE

Maya Chen

Maya Chen

Lead AI Researcher

Published: July 16, 2026
Hy3 open-source mixture-of-experts model reference card

TLDRHy3 launched July 6, 2026 as Tencent Hunyuan's open-source 295B MoE (21B active, 256K context) at $0.14/$0.58 per M tokens. Apache 2.0 weights on Hugging Face.

What Is Hy3? Tencent's 295B MoE That Topped OpenRouter at $0.58 per Million Output Tokens

Hy3 is a 295-billion-parameter open-source mixture-of-experts language model from Tencent Hunyuan, launched July 6, 2026 under Apache 2.0. It activates about 21B parameters per token, ships with a 256K context window, and lists official API pricing at 1 RMB per million input tokens and 4 RMB per million output tokens (roughly $0.15 and $0.59 USD). Weights are published on Hugging Face, GitHub, ModelScope, and GitCode, and the model topped OpenRouter's weekly usage leaderboard shortly after launch.

Key Takeaways

  • Hy3 is a 295B total / 21B active MoE with a 3.8B multi-token prediction (MTP) layer, 192 experts, and top-8 routing.
  • The full open-source release landed July 6, 2026 under Apache 2.0, following the Hy3-preview from late April 2026 that used a more restrictive license.
  • Official API pricing is 1 RMB input / 4 RMB output per million tokens; OpenRouter mirrors this at $0.14 / $0.58 with a free tier running through July 21, 2026.
  • Tencent's blind expert evaluation over 270 real-world workflows scored Hy3 at 2.67/4 vs GLM-5.1 at 2.51/4, per the official launch blog.
  • Official 1-bit and 4-bit GGUF quantizations shipped July 14, 2026, letting the model run on a single ~96GB GPU via llama.cpp with MTP.
  • Community sentiment centers on coding, tool-calling, and 256K-context agent workloads, where early testers describe Hy3 as "practical over hype."

What Is Hy3?

Hy3 is the third-generation flagship large language model from Tencent's Hunyuan research team. It is a sparse mixture-of-experts (MoE) architecture: the model contains 295 billion total parameters but only activates around 21 billion per token via top-8 routing across 192 experts. A dedicated 3.8B multi-token prediction (MTP) layer sits on top, letting the model draft several tokens ahead per forward pass — a design choice that shows up prominently in Tencent's speed claims.

The model is aimed squarely at the workloads where 2026 open-weight LLMs have been converging: long-context coding, tool-calling agents, and cost-sensitive production inference. Tencent positions Hy3 as a general-purpose model with particular strength in frontend generation, CI/CD workflows, and data/storage tasks, based on the internal blind evaluation reported in the official Hy3 announcement thread.

The Hy3-preview release surfaced in late April 2026 with the same 295B / 21B-active shape but a more restrictive license. Commenters on the r/LocalLLaMA preview thread described that preview as "barely even open-weights," which the July 6 Apache 2.0 release effectively answered.

Hy3 at a Glance

AttributeValue
DeveloperTencent Hunyuan
TypeMixture-of-Experts (MoE) large language model
Total parameters295B
Active parameters~21B per token
Extra components3.8B multi-token prediction (MTP) layer
Experts / routing192 experts, top-8
AttentionGQA (64 heads / 8 KV)
ModalityText in, text out
Context window256K tokens
LicenseApache 2.0 (full release, July 6, 2026)
Official API pricing1 RMB / 4 RMB per M tokens (input / output); 0.25 RMB cached input
USD equivalent~$0.15 / $0.59 per M tokens
WeightsHugging Face, GitHub, ModelScope, GitCode
QuantizationsOfficial 1-bit and 4-bit GGUF, GPTQ Int4 (via AngelSlim, July 14, 2026)
Artificial Analysis Intelligence Index41
SWE-bench Verified78 resolved (per Hugging Face card)
AvailabilityGA via official API and OpenRouter (launched July 6, 2026)

How Hy3 Works and What Makes It Different

Hy3's architecture is what Tencent calls a Sparse-21B-Active MoE: 192 experts with top-8 routing keep the per-token compute footprint near a dense 21B model, while the full 295B parameter pool provides capacity for specialization. This is the same broad recipe used by other 2026 open-weight releases such as DeepSeek V4, which we covered in our DeepSeek V4 release page.

Three details set Hy3 apart in the current landscape:

  • MTP-Accelerated Decoding. The 3.8B multi-token prediction head lets Hy3 draft multiple tokens per step, which Tencent uses to sustain higher throughput at long contexts. Community tests report that MTP gains fade at very long context lengths, according to Artificial Analysis notes and independent runs.
  • Blind-Panel Evaluation. Rather than only leaning on public leaderboards, Tencent's launch blog reports a blind expert panel over 270 real-world workflows, scoring Hy3 at 2.67/4 versus GLM-5.1 at 2.51/4. The same evaluation reports hallucination rate dropping from 12.5% in the preview to 5.4%, and multi-turn issue rate from 17.4% to 7.9%.
  • AngelSlim Quantization Pipeline. Tencent shipped official 1-bit and 4-bit GGUF quantizations plus GPTQ Int4 within eight days of launch, targeting single-GPU llama.cpp deployment with the MTP path preserved, per the July 14 announcement.

One anchor quote from the community frames the reception well: "It's performing about on-par with GLM 5.2 for my needs (subagent orchestration mostly). It does use a lot of tokens so not quite as cheap as it looks, but still very cost effective," wrote one user on the r/AIToolsPerformance discussion.

What You Can Do With Hy3

The workloads showing up most in early adoption are the ones Hy3 was tuned for:

  • Long-context coding and repo navigation. The 256K window plus SWE-bench Verified score of 78 resolved (per the Hugging Face card) put Hy3 in the practical range for agentic coding tasks. Kilo Code integrated Hy3 into its coding agent at launch as tencent/hy3:free.
  • Subagent orchestration. Community reports on OpenRouter specifically call out subagent workflows where Hy3's output pricing makes multi-call plans affordable.
  • Tool-calling agents. Hermes agents and other frameworks made Hy3 free for two weeks after launch, citing strong tool-calling reliability at 256K context.
  • On-prem inference on a single high-memory GPU. The 1-bit GGUF build targets one ~96GB GPU with llama.cpp; the 4-bit build is claimed to be near-full quality, though sustained long-context multi-user loads remain to be independently verified.
  • Cost-sensitive production replacement of mid-tier closed models. At $0.58 per million output tokens, Hy3 sits about 17x cheaper than Claude Sonnet 5's $10 output rate, based on OpenRouter listings referenced in the community discussion.

How Hy3 Compares

The relevant comparison for Hy3 is not the frontier closed tier — it is the crowded mid-tier of open-weight and low-cost models where price-per-quality actually decides which model wins the router.

ModelParams (active)ContextOutput $/MLicense
Hy3295B (21B) MoE256K$0.58Apache 2.0
GLM 5.2Not yet confirmed1M$2.86Proprietary
Claude Sonnet 5Not yet confirmedNot yet confirmed$10.00Proprietary

For a closed-model reference point on the higher end of that pricing spread, see our Claude Sonnet 5 deep dive. For an open-weight peer at a very different scale, our Qwen 3.6 27B analysis is a useful contrast.

Availability: How to Access Hy3

Hy3 is generally available through multiple official channels:

  • Official Tencent Hunyuan API. The Hy blog lists 1 RMB per million input tokens, 4 RMB per million output tokens, and 0.25 RMB per million cached input tokens as the standard rate.
  • Open weights on Hugging Face. The tencent/Hy3 repository ships full-precision weights, model card, and reference evals under Apache 2.0.
  • Mirror hosts. Weights are additionally distributed on GitHub, ModelScope, and GitCode per the launch blog.
  • Quantized builds. Official 1-bit and 4-bit GGUF plus GPTQ Int4 quantizations are published via AngelSlim, released July 14, 2026.
  • Router access. OpenRouter carries Hy3 at $0.14 / $0.58 per million tokens, with a free tier ID tencent/hy3:free available through July 21, 2026. Artificial Analysis tracks the model's intelligence index at 41 and provides ongoing throughput measurements.

What We Don't Know Yet

A few questions remain open even after the full release:

  • Sustained quality of the 1-bit quantization on multi-user, long-context serving is not yet independently benchmarked.
  • Full independent leaderboards against the very latest DeepSeek, GLM, Claude, and GPT variants are still sparse; official evals lean on Tencent's blind panel.
  • How pricing, rate limits, and throughput behave after OpenRouter's free tier ends on July 21, 2026 is not yet visible.
  • The commercial friction of self-hosting or fine-tuning a 295B-class MoE for downstream users has not been broadly reported.

Frequently Asked Questions

What is Hy3?

Hy3 is Tencent Hunyuan's third-generation large language model, launched July 6, 2026 as an Apache 2.0 open-source mixture-of-experts with 295B total parameters and around 21B active per token. It supports a 256K context window and is available via Tencent's official API, Hugging Face weights, and OpenRouter.

Who makes Hy3?

Hy3 is developed by the Hunyuan (Hy) research team inside Tencent. Announcements come through the official Tencent Hunyuan blog and the @TencentHunyuan account, with weights distributed on Hugging Face, GitHub, ModelScope, and GitCode.

How much does Hy3 cost?

Hy3 costs 1 RMB per million input tokens and 4 RMB per million output tokens on Tencent's official API, with cached input at 0.25 RMB per million. On OpenRouter that maps to $0.14 input and $0.58 output per million tokens, and a free tier (tencent/hy3:free) is available through July 21, 2026.

Is Hy3 open source?

Yes, Hy3 is released under Apache 2.0 as of July 6, 2026, with full weights on Hugging Face. The earlier Hy3-preview from late April 2026 used a more restrictive license that community members characterized as "weights available" rather than fully open.

What is Hy3's context window?

Hy3 supports 256K tokens of context. That is shorter than the 1M-token windows some closed 2026 models advertise but long enough for most agentic coding and multi-document reasoning workloads.

How does Hy3 compare to GLM 5.2?

Hy3 is roughly 5x cheaper than GLM 5.2 on output tokens ($0.58 vs $2.86 per million on OpenRouter). Community reports place Hy3's quality between GLM 5.1 and GLM 5.2 on subagent and coding tasks, though independent apples-to-apples benchmarks are still limited.

Can Hy3 run locally?

Yes. Tencent released official 1-bit and 4-bit GGUF quantizations plus GPTQ Int4 weights on July 14, 2026 via AngelSlim. The 1-bit build runs on a single ~96GB GPU with llama.cpp and multi-token prediction (MTP) support.

What to Watch Next

Three signals will shape how this page evolves. First, what happens to Hy3's OpenRouter usage and pricing after the free tier expires on July 21, 2026. Second, independent benchmark results — Artificial Analysis's intelligence index and community coding evals — as they accumulate against DeepSeek V4, GLM 5.2, and the newer closed frontier. Third, real quality measurements from the 1-bit GGUF build in multi-user production, which will decide whether "single-GPU 295B" is a serving reality or a demo.

Building similar long-context chat and coding workloads? On kie.ai you can try Claude Opus 5, Gemini 3 Pro, and GPT-5.6.

Maya Chen

About Maya Chen

Maya tracks AI model releases, benchmarks, and developer adoption signals across the open and closed model landscape.

View all posts by Maya Chen