What Is Tencent Hy4? 770B MoE, 1M Context

Priya Nair

Priya Nair

AI Infrastructure Analyst

Published: August 30, 2026
Tencent Hy4 Preview model overview

TLDRTencent Hy4 has 770B total parameters, 49B active, and 1M+ context for coding, office work, game development, and research.

What Is Tencent Hy4? The 1M-Token Open MoE

Tencent Hy4 is Tencent’s open-source preview Mixture-of-Experts language model for coding, productivity, game development, and scientific research. Tencent released Hy4 Preview on August 28, 2026, with 770B total parameters, 49B active parameters per token, and a context window exceeding 1M tokens. It is available through Tencent products, Tencent Cloud TokenHub, and released model weights, but its final release status and official pricing are not yet confirmed.

Key Takeaways

  • Tencent Hy4 Preview is a general-purpose language model developed by the Tencent Hy Team.
  • The model contains 770B total parameters but activates 49B parameters for each token.
  • Its context window exceeds 1M tokens, targeting large repositories, documents, and long-running agents.
  • The model uses Apache License 2.0, according to its official Hugging Face model card.
  • Tencent reports a 2.99/4.00 average in an internal blind test against 2.92 for GLM-5.3 and 2.94 for Kimi K3.
  • Early community tests suggest strong coding and 3D-generation ability, but cost and latency results vary substantially.

What Is Tencent Hy4?

Tencent Hy4 Preview is the fourth-generation model in Tencent’s Hy family and the company’s new open-weight flagship. Its formal model-card name is tencent/Hy4-preview, and the model is described as a new-generation Mixture-of-Experts, or MoE, system.

Tencent released the preview for real-world productivity tasks rather than positioning it only as a conversational chatbot. The stated target workloads include software engineering, office analysis, game development, finance, security, and scientific research. The model’s 770B total parameter count gives it a large expert pool, while the 49B active count reduces the amount of computation selected for each token.

Tencent Hy4 Preview is a 770B-parameter MoE model that activates 49B parameters per token and supports a context window exceeding 1M tokens.

The model was publicly announced on August 28, 2026, through Tencent’s Hy4 announcement. The word “Preview” matters: the available checkpoint is an early version, not evidence of a final Hy4 release. Tencent’s materials describe additional room for improvement in both pre-training and post-training.

Tencent Hy4 at a Glance

SpecificationTencent Hy4 Preview
DeveloperTencent Hy Team
TypeLarge language model; Mixture-of-Experts
ModalityText generation
Total parameters770B
Active parameters49B per token
Context windowExceeding 1M tokens
PricingNot yet confirmed; one third-party report lists $0.83 per 1M input tokens and $2.50 per 1M output tokens
AvailabilityTencent products, Tencent Cloud TokenHub, and released weights
LicenseApache License 2.0
Release statePreview model released on August 28, 2026

The official Hy4 Preview model card identifies the model as a Tencent Hy Team release and lists Apache 2.0 licensing. The model card also documents deployment paths through common inference frameworks.

Tencent released Hy4 Preview, a new open weight model under Apache License 2.0! > Hy4 preview is

Source: @testingcatalog

How Tencent Hy4 Works and What Makes It Different

Sparse Expert Routing

Tencent Hy4’s main architectural anchor is Sparse Expert Routing. The model has 78 backbone layers. The first layer uses a standard dense feed-forward network, while the remaining 77 layers use MoE blocks containing 256 routed experts and one shared expert. Each token selects the top 8 routed experts alongside the shared expert.

This structure separates model capacity from per-token compute. The full checkpoint still contains 770B parameters, so storage and serving remain demanding. However, only 49B parameters are active for an individual token, allowing the model to use a much larger learned capacity than its active compute figure suggests.

A day-one note from the vLLM project reports additional implementation details, including 21 layers that compute their own sparse index. It also says each query attends to 2,048 tokens in the relevant sparse-attention implementation. That means a 1M-token context window should not be interpreted as dense, full-attention processing over 1M tokens at every operation.

1M Context

1M Context is the second important Hy4 anchor. A context window exceeding 1M tokens can hold a large software repository, extensive logs, many business documents, or a long sequence of tool calls in one session.

Context capacity does not guarantee perfect retrieval or reasoning. Long inputs can contain irrelevant material, conflicting instructions, or repeated evidence. Engineering teams should test repository recall, instruction following, tool-call consistency, and output quality at different context lengths rather than treating the headline number as a quality score.

Native Multi-Token Prediction

The architecture also includes Native Multi-Token Prediction. According to vLLM’s implementation notes, the checkpoint includes a 10B multi-token prediction layer with 0.7B active parameters and draft depth 3.

This mechanism is intended to support speculative decoding. A draft component proposes multiple future tokens, which can improve serving efficiency when the inference stack accepts those predictions. Actual throughput depends on hardware, quantization, batch size, context length, and framework implementation.

Agentic Research

Tencent presents Agentic Research as a major capability direction. The model participated in workflows that propose approaches, run experiments, inspect results, and iterate on subsequent work. Tencent says Hy4 helped optimize training methods, data strategies, evaluation frameworks, and low-level operators during its own development process.

That claim describes a vendor-reported development workflow, not autonomous general intelligence. It shows how Tencent is using Hy4 within engineering loops, while independent replication of the complete process remains unavailable in the signal bundle.

What You Can Do With Tencent Hy4

Tencent designed Hy4 for tasks that require planning across multiple steps. Software engineering is the clearest use case. The model is intended to understand requirements, plan changes, debug failures, validate results, and work across long-context projects.

Office and analytical work is another target. Tencent describes workflows that move from information processing to documents, spreadsheets, presentations, and financial analysis. The model can therefore be evaluated on cross-document synthesis, structured output, and analysis involving messy source material.

Game development is also part of the model’s stated scope. Tencent says Hy4 can generate a playable prototype from a natural-language request and continue refining a complex project through multi-turn interaction. Early community demonstrations reported browser-based games, 3D scenes, and interactive visual prototypes. Those demonstrations are useful capability signals, but they are not standardized benchmarks.

Scientific workloads include AI research, molecular dynamics, condensed-matter physics, and fundamental mathematics. These tasks require reproducible calculations and source checking. Hy4 can assist with exploration and implementation, but its outputs still need domain-expert review.

How Tencent Hy4 Compares

Tencent reports a blind internal evaluation involving 163 experts and 203 engineering tasks. Hy4 Preview averaged 2.99 out of 4.00, compared with 2.92 for GLM-5.3 and 2.94 for Kimi K3.

Internal blind evaluationTencent Hy4 PreviewComparison result
Versus GLM-5.32.99/4.00GLM-5.3: 2.92/4.00
Versus Kimi K32.99/4.00Kimi K3: 2.94/4.00
Evaluation scale4.00 pointsNot an independent leaderboard

On Tencent’s internal blind test, Hy4 Preview averaged 2.99/4.00 against 2.92 for GLM-5.3 and 2.94 for Kimi K3; independent confirmation remains unconfirmed.

A third-party summary of Tencent’s reported evaluations gives Hy4 a 85.4 score on Terminal-Bench 2.1 and 64.3 on DeepSWE, compared with 28.0 for Hy3 on the latter measure. Those figures are not independently verified in the available material. Another community report places Hy4 at 1,633 points and approximately fifth overall in Code Arena WebDev, around 115 points above Hy3.

For context on nearby releases, see the earlier analysis of GLM-5.3 and the coverage of Kimi K3’s release.

Availability: How to Access Tencent Hy4

Tencent Hy4 Preview is available through WorkBuddy, CodeBuddy, Yuanbao, and ima. Tencent stated that Hy4 Preview would be free on WorkBuddy and CodeBuddy for 2 weeks after launch. That introductory access period does not establish permanent free availability.

API access is available through Tencent Cloud TokenHub, according to Tencent’s release materials. The official model weights are also published through Tencent’s Hy4 Preview GitHub repository and Hugging Face. The model card documents serving with vLLM and SGLang, along with compatible local deployment paths.

The 770B checkpoint is not a typical laptop model. A single third-party report claims FP8 serving requires tensor parallelism across 8 devices, but the complete hardware and throughput profile remains unconfirmed. Teams considering self-hosting should wait for validated memory, quantization, throughput, and licensing guidance for their deployment configuration.

For teams evaluating adjacent long-context coding models while Hy4 infrastructure is assessed, Kimi K3 is a separate model option, not a Tencent Hy4 access route. Kie.ai does not currently host Tencent Hy4.

What We Don’t Know Yet

Several practical questions remain open:

  • Tencent has not published a final Hy4 release date or confirmed whether Preview behavior will remain unchanged.
  • Official API pricing, cache pricing, rate limits, and regional availability are not confirmed in the release materials.
  • One report lists $0.83 per 1M input tokens and $2.50 per 1M output tokens, while community reports describe widely different task-level costs.
  • Independent testing has not established whether the reported Terminal-Bench, DeepSWE, Code Arena, or internal blind-evaluation results generalize across workloads.
  • The practical quality of the 1M-token context remains unclear, especially for retrieval, sustained tool use, and long sessions.
  • Hardware requirements, quantization choices, serving throughput, and production latency need broader reproducible testing.
  • Early reports describe excessive reasoning time and a tendency to over-verify work; the frequency and severity of these behaviors are not yet established.

Frequently Asked Questions

What is Tencent Hy4?

Tencent Hy4 is Tencent’s open-source preview Mixture-of-Experts language model for coding, productivity, game development, and scientific research. It has 770B total parameters, activates 49B parameters per token, and supports a context window exceeding 1M tokens.

Is Tencent Hy4 open source?

Tencent Hy4 Preview is released as an open-source, open-weight model under the Apache License 2.0. Its model weights and model documentation are available through Tencent’s official repositories.

How many parameters does Tencent Hy4 have?

Tencent Hy4 has 770B total parameters and 49B active parameters per token. Its Mixture-of-Experts design activates only a subset of experts during each token calculation.

What is Tencent Hy4’s context window?

Tencent Hy4 has a context window exceeding 1M tokens. That capacity is intended for large codebases, document collections, logs, and extended agent sessions.

How much does Tencent Hy4 cost?

Tencent Hy4’s official price is not yet confirmed in the available release materials. One third-party report lists $0.83 per 1M input tokens and $2.50 per 1M output tokens, but those figures remain unconfirmed.

How can I access Tencent Hy4?

Tencent Hy4 Preview can be accessed through Tencent products including WorkBuddy, CodeBuddy, Yuanbao, and ima, and through Tencent Cloud TokenHub for API access. The released weights are also available through Tencent’s official model repositories for compatible self-hosting.

Tencent Hy4 vs GLM-5.3: which is better?

Tencent Hy4 scored slightly above GLM-5.3 in Tencent’s internal blind evaluation, averaging 2.99 out of 4.00 versus 2.92. That result is narrow, internally produced, and does not establish that Hy4 is better for every workload.

What to Watch Next

The most important updates are independent long-context tests, verified API pricing and latency, and clearer hardware guidance for serving the 770B checkpoint. A final Hy4 release, broader third-party benchmarks, or changes to the preview’s known reasoning behavior would materially update this reference page.

Building similar 1M-token open MoE coding and research workflows? On kie.ai you can try DeepSeek-V4.1-Flash, Kimi K3, and Claude Opus 5.5.

Priya Nair

About Priya Nair

Priya covers serving costs, context windows, and the infrastructure tradeoffs behind each model launch.

View all posts by Priya Nair