Head-to-Head: Tencent Hy4 and GLM-5.3 on 1M-Token Context

Maya Chen

Maya Chen

Lead AI Researcher

Published: September 6, 2026
Tencent Hy4 and GLM-5.3 comparison

TLDRA 1M-token context and 2.99/4.00 internal score give Tencent Hy4 an early edge over GLM-5.3, but independent head-to-head data remains limited.

Tencent Hy4 has the documented edge in Tencent’s internal engineering evaluation and offers open weights with a 1M-token context, while GLM-5.3 is supported here mainly by a narrow vendor-reported comparison; the right choice depends on deployment control, workload cost, and independent validation.

Key Takeaways

  • Tencent Hy4 Preview is a 770B-parameter Mixture-of-Experts model with 49B active parameters per token and a 1M-token context window.
  • Tencent’s internal blind evaluation scored Hy4 at 2.99 out of 4.00, compared with 2.92 for GLM-5.3 across 203 engineering tasks judged by 163 experts.
  • The score edge is narrow and vendor-reported. Independent, reproducible GLM-5.3 versus Hy4 testing remains limited.
  • Hy4 is released under the Apache License 2.0 with model weights, deployment documentation, and an FP8 variant. GLM-5.3 licensing and weight availability are unverified in this bundle.
  • A third-party listing reported Hy4 pricing at $0.83 per million input tokens and $2.50 per million output tokens. A community comparison reported GLM-5.3 at $1.40 input and $4.40 output, but neither schedule is established as an official long-term price.
  • Choose Hy4 for self-hosting, very long repositories, and experiments with open-weight agent workflows. Choose GLM-5.3 only when its specific endpoint, pricing, and reliability have been verified for the intended workload.

Tencent Hy4 vs GLM-5.3 at a Glance

DimensionTencent Hy4 PreviewGLM-5.3
Model architecture770B total parameters; 49B active per token; sparse MoENot yet confirmed
Context window1M tokensNot yet confirmed
Licensing and weightsOpen weights; Apache License 2.0Not yet confirmed
Direct evaluationTencent internal blind test: 2.99/4.00 across 203 engineering tasks2.92/4.00 in the same Tencent-reported comparison
Reported API pricingCommunity and third-party reports: $0.83/M input; $2.50/M outputCommunity report: $1.40/M input; $4.40/M output
Deployment accessTencent products, Tencent Cloud TokenHub, Hugging Face, GitHub, and compatible inference stacksNot yet confirmed

The table is intentionally asymmetric. Hy4 has an official model card, a public repository, and a Tencent release statement in the supplied record. GLM-5.3 appears in the comparison evidence, but the bundle does not provide a matching official model card or full technical specification.

Capabilities and Architecture

Tencent Hy4 Preview is designed as a productivity-oriented language model rather than a general specification exercise. Tencent positions it for software engineering, office analysis, game development, finance, security, and scientific research. Its stated software-engineering focus includes planning, debugging, validation, and long-horizon development tasks.

The official Hy4 Preview repository documents a 78-layer backbone. The first layer uses a dense feed-forward network. The remaining 77 layers use Mixture-of-Experts blocks containing 256 routed experts and 1 shared expert. Each token activates 8 routed experts plus the shared expert.

Tencent released Hy4 Preview, a new open weight model under Apache License 2.0! > Hy4 preview is

Source: @testingcatalog

Hy4’s total parameter count is 770B, but its active parameter count is 49B per token. That distinction matters for inference planning. The full checkpoint remains demanding to store and serve, even though sparse routing limits per-token computation. The model also includes a 10B-parameter native multi-token prediction layer with 0.7B active parameters for speculative decoding.

Its attention design is called Gated DeepSeek Sparse Attention, or Gated DSA. IndexCache reuses sparse attention indices across layers, a design intended to reduce repeated work in long contexts. The repository lists a 6,144-dimensional hidden size, 64 attention heads, 2,048-token query compression, and 512-dimensional key-value compression.

The 1M-token context is the clearest technical difference documented in this bundle. It can accommodate large repositories, extensive logs, or multiple office documents in one workflow. It does not guarantee perfect recall. Long-context systems still require retrieval checks, instruction testing, and context-budget monitoring.

GLM-5.3 has no equivalent architecture record in the supplied evidence. The bundle does mention GLM-5.3 in Tencent’s internal evaluation and mentions GLM-5.3 Flash in public comparison pages, but those references should not be treated as a complete specification for GLM-5.3.

A separate analysis of Tencent Hy4’s 1M-token open MoE design covers the model’s deployment implications in more detail. The key comparison point is simple: Hy4’s architecture, license, context window, and model files are documented here; GLM-5.3’s corresponding cells remain unverified.

Benchmarks and Evidence Quality

The strongest direct comparison is Tencent’s internal blind evaluation. Tencent says 163 experts assessed outputs on 203 engineering tasks. Hy4 Preview averaged 2.99 out of 4.00, compared with 2.92 for GLM-5.3 and 2.94 for Kimi K3.

Tencent’s reported score gives Hy4 a 0.07-point advantage over GLM-5.3 on that evaluation. That is a useful signal, not proof of broad superiority. The test was conducted internally, and the bundle does not provide a fully reproducible protocol, task distribution, prompts, or independent replication.

Tencent’s 2.99 versus GLM-5.3’s 2.92 is evidence of a narrow internal-test lead, not a universal model ranking.

Tencent also reports 85.4 on Terminal-Bench 2.1 and 64.3 on DeepSWE for Hy4 Preview. Those figures are relevant to coding-agent discussions, but the corresponding GLM-5.3 scores are not supplied. A score for only one model describes coverage, not a head-to-head win.

Community summaries repeat a reported Hy4 win rate of 46.8%, with 12.8% ties and 40.4% losses against GLM-5.3 in the internal comparison. Those percentages come from community reposts of launch materials rather than an independent study. They also show why “slightly ahead” is the appropriate reading: Hy4 did not win every comparison.

There is one additional public comparison that must be kept separate. BenchLM reports GLM-5.3 Flash at 66.09 and Hy4 Preview at 65.94, with overlapping 90% intervals. Its page labels Hy4’s score as estimated and compares GLM-5.3 Flash, not the base GLM-5.3 target in this article. It therefore cannot settle the Tencent Hy4 versus GLM-5.3 question.

Building similar long-context model comparisons? On kie.ai you can try DeepSeek-V4.1-Flash [Chat] <!-- rel:dofollow class:self -->, Kimi K3 [Chat] <!-- rel:dofollow class:self -->, and Claude Opus 5.5 [Chat] <!-- rel:dofollow class:self -->.

Pricing, Deployment, and Limits

Pricing evidence is less reliable than the model specifications. A third-party catalog reported Tencent Hy4 at $0.83 per million input tokens and $2.50 per million output tokens. A Reddit comparison reported GLM-5.3 at $1.40 per million input tokens and $4.40 per million output tokens.

Those figures suggest Hy4 may be less expensive on both token categories in that particular comparison. They are not enough to establish a permanent price advantage. The bundle does not provide a single official pricing schedule for either model, and hosted prices can vary by region, provider, caching policy, or preview status.

The available price reports favor Hy4, but the pricing comparison remains community-sourced and should be rechecked before budgeting.

Early community testing also reports inconsistent task economics. One same-prompt 3D task reportedly took Hy4 more than 2 hours and cost $5.40, while another finished in 10 minutes and cost $1.30. A separate tester estimated Hy4 at 8 times the cost and 4 times the latency of Hy3 on one prompt. These are not GLM-5.3 comparisons, but they show why output-token price alone is not a sufficient operating-cost metric.

Hy4’s official access paths are clearer. Tencent says the model is available through WorkBuddy, CodeBuddy, Yuanbao, ima, Tencent Cloud TokenHub, and released weights. The Hugging Face model card lists Transformers, vLLM, and SGLang deployment paths. Tencent also released an FP8 variant and quantization tooling.

The full model is not a lightweight local deployment. A community quantization report described compression from roughly 1.5TB to about 214GB using a low-bit format, with reported score changes on MCP Atlas and SWE-Bench Multilingual. That report is an early community result, not a guaranteed hardware recipe or throughput benchmark.

For a separate hosted chat baseline while checking model access, developers can inspect Kimi K3. It should be treated as a different model and workflow reference, not as evidence about GLM-5.3 or Tencent Hy4.

Hy4 also has acknowledged limitations. Tencent’s release materials describe longer-than-necessary reasoning on complex tasks and a tendency to over-verify its own work. Early users similarly report long agent loops, high token consumption, and variable latency. These issues matter for interactive coding, where a lower token price can be offset by repeated tool calls.

Which One Should You Use?

Choose Tencent Hy4 if:

  • The project requires open weights, Apache 2.0 licensing, or control over the serving stack.
  • The workload includes repositories, document collections, or logs approaching a 1M-token context.
  • The team is building long-horizon coding, research, game-development, or multi-file office agents.
  • The team can operate multi-GPU infrastructure or use a supported hosted channel.
  • A vendor-reported 2.99 out of 4.00 engineering score is a useful starting signal, followed by local validation.

Choose GLM-5.3 if:

  • A verified GLM-5.3 endpoint already exists in the team’s preferred infrastructure.
  • The workload has a tested GLM-5.3 prompt, tool schema, and cost profile.
  • The team prefers to wait for independent evaluations that compare identical tasks.
  • A GLM-5.3 deployment offers better measured latency, availability, or integration for the specific application.

The responsible current verdict is conditional: Hy4 is the better-documented option for open deployment and very long context, while GLM-5.3 cannot be ruled out on quality or operating experience because comparable independent evidence is still missing.

Frequently Asked Questions

Is Tencent Hy4 better than GLM-5.3?

Tencent Hy4 has the early documented advantage in Tencent’s internal engineering evaluation, scoring 2.99 out of 4.00 versus 2.92 for GLM-5.3. That result is vendor-reported, and independent head-to-head evidence is not yet sufficient to name Hy4 the universal winner.

Which is better for coding, Tencent Hy4 or GLM-5.3?

Tencent Hy4 is the better-documented coding choice because Tencent specifically targets long-horizon software engineering and reports 85.4 on Terminal-Bench 2.1. A directly matched independent coding score for GLM-5.3 is not confirmed in the supplied evidence.

Is Tencent Hy4 cheaper than GLM-5.3?

A third-party listing reported Tencent Hy4 at $0.83 per million input tokens and $2.50 per million output tokens. A community comparison reported GLM-5.3 at $1.40 per million input tokens and $4.40 per million output tokens, but neither figure is confirmed as a stable official schedule.

Does Tencent Hy4 have a larger context window than GLM-5.3?

Tencent Hy4 has a confirmed context window of 1 million tokens. The supplied evidence does not confirm the GLM-5.3 context size, so a direct context-window advantage cannot be calculated.

Is Tencent Hy4 open source while GLM-5.3 is not?

Tencent Hy4 is released with open weights under the Apache License 2.0. GLM-5.3’s license and weight-availability status are not confirmed in this bundle, so the comparison should not assert that GLM-5.3 is closed.

How can developers access Tencent Hy4 and GLM-5.3?

Developers can access Tencent Hy4 through Tencent products, Tencent Cloud TokenHub, and the released model weights using documented inference stacks. A confirmed official GLM-5.3 access path is not documented in the supplied comparison evidence.

What to Watch Next

The most useful updates will be an independent GLM-5.3 versus Hy4 coding evaluation, stable official pricing for both models, and measured throughput across common multi-GPU configurations. Long-context recall, tool-call reliability, and agent-loop token usage also deserve controlled testing.

Building similar long-context coding and agent workflows? On kie.ai you can try Claude Opus 5, OpenAI Codex, and Kimi K3.

Maya Chen

About Maya Chen

Maya tracks AI model releases, benchmarks, and developer adoption signals across the open and closed model landscape.

View all posts by Maya Chen