What Is GLM-5.3 Flash? 1M-Token Multimodal MoE

Maya Chen

Maya Chen

Lead AI Researcher

Published: August 30, 2026
GLM-5.3 Flash model overview

TLDR320B total parameters, 18B active, and a 1,048,576-token context define GLM-5.3 Flash, Z.ai's open-weight multimodal model.

GLM-5.3 Flash 101: 1M-Token Multimodal MoE

GLM-5.3 Flash is an open-weight, natively multimodal mixture-of-experts model developed by Z.ai for coding, agentic reasoning, and visual software workflows. It combines 320B total parameters with 18B active parameters, supports a 1,048,576-token context window, and accepts text, images, video, and files. Z.ai released it on August 26, 2026, after testing the model anonymously under the codename Ox Alpha.

GLM-5.3 Flash is designed to deliver much of a large model's capability with lower active computation and lower serving cost. Its weights use the MIT License, and its documented maximum output is 128,000 tokens.

Updated 2026-09-14: NVIDIA has published an official NVFP4 quantization of GLM-5.3 Flash (see the Update below).

Key Takeaways

  • GLM-5.3 Flash is Z.ai's first natively multimodal model in the GLM-5 series.
  • Its architecture has 320B total parameters and 18B active parameters per token.
  • The context window contains 1,048,576 tokens, or approximately 1 million tokens.
  • The model supports coding agents, function calling, browser and GUI workflows, Blender tasks, and document production.
  • Cloudflare's model listing gives a reference price of $0.15 per million input tokens and $0.50 per million output tokens.
  • The official weights are public, but practical local inference generally requires quantization and substantial memory.

What Is GLM-5.3 Flash?

GLM-5.3 Flash is the efficiency-focused member of Z.ai's GLM-5.3 family. It is not simply a smaller dense language model. It is a sparse mixture-of-experts system that contains 320B parameters in total but activates 18B parameters for each token.

The model is aimed at workloads where reasoning quality, tool use, and operating cost matter together. Those workloads include repository-level coding, long-running agents, visual frontend development, browser interaction, and structured professional work.

Z.ai describes GLM-5.3 Flash as the first native multimodal model in the GLM-5 series. Its inputs include video, image, text, and files, while its output modality is text. Function calling, reasoning, streaming output, structured output, and context caching are part of the documented capability set.

The model first appeared anonymously as Ox Alpha on OpenCode and OpenRouter before its public identification. Z.ai says the anonymous test gathered user feedback before release. An earlier technical backgrounder covers the model's Ox Alpha phase in more detail: Meet Ox Alpha, the Free Stealth Model With a 1M-Token Context Window.

GLM-5.3 Flash at a Glance

SpecificationConfirmed information
DeveloperZ.ai
Release statusReleased on August 26, 2026
TypeOpen-weight sparse mixture-of-experts language model
Total parameters320B
Active parameters18B
Input modalityVideo, image, text, and file
Output modalityText
Context window1,048,576 tokens
Maximum output128,000 tokens
Pricing$0.15 per 1M input tokens and $0.50 per 1M output tokens on Cloudflare Workers AI
AvailabilityZ.ai API, GLM Coding Plan, ZCode, chat products, and public weights
LicenseMIT License

Cloudflare documents the model as @cf/zai-org/glm-5.3-flash, with paid access required for its Workers AI deployment. Its listing also identifies reasoning, vision, and function calling support: Cloudflare's GLM-5.3 Flash documentation.

How GLM-5.3 Flash Works and What Makes It Different

The central design choice is Hybrid Sparse-Linear Attention. Linear attention handles local dependencies through state modeling, while sparse attention retrieves relevant global context through an indexed mechanism. This reduces the cost of processing long prompts without removing access to distant information.

Z.ai reports that the design reduces attention computation by 3.01 times and KV-cache size by 4.44 times compared with GLM-5.3. These are architectural comparisons from the developer's release material, not independent serving measurements.

A related component called IndexPool compresses four indexer key vectors into one through weighted pooling. The goal is to reduce indexer latency and memory overhead when the model operates near its 1-million-token context limit.

The model also uses Manifold-Constrained Hyper-Connections, or mHC, to improve scaling efficiency. Z.ai pairs those changes with a 30T-token multimodal pretraining corpus.

The sparse design explains the difference between total and active parameters. The model stores 320B parameters, but only 18B are activated for an individual token. That can lower per-token computation while preserving a much larger pool of learned experts.

GLM-5.3 Flash is a 320B-parameter MoE with 18B active parameters and a 1,048,576-token context window.

The model's visual capability is integrated into the coding loop rather than treated as a separate image question-answering feature. Z.ai calls this Native Multimodal Visual Coding. The model can inspect interfaces, rendered results, and interaction feedback while coordinating work across code, browsers, and graphical interfaces.

What You Can Do With GLM-5.3 Flash

GLM-5.3 Flash is primarily relevant to engineering teams building agents that must act, inspect results, and revise work.

Typical uses include:

  • Writing, debugging, refactoring, and explaining software.
  • Operating coding agents across repositories and multi-step tasks.
  • Generating frontend applications and checking rendered interfaces.
  • Controlling browsers and desktop interfaces through computer-use workflows.
  • Creating Blender scenes and other structured 3D outputs through tool integrations.
  • Producing PPTX, PDF, DOCX, and XLSX deliverables from research or office tasks.
  • Processing documents, conducting financial research, and summarizing files.
  • Writing prose and creative material, although community speed tests indicate that prose generation can be slower than coding on some local setups.

A public Blender/MCP test reported 811 objects in 38 minutes and 52 seconds at a stated cost of $0.0526, compared with 847 objects in 40 minutes and 43 seconds at $0.8807 for GLM-5.3. That result is a single community test, not a universal cost guarantee: the Blender comparison report.

GLM-5.3 Flash is a practical candidate for multimodal coding and tool-use workloads where cost matters, but throughput depends heavily on the runtime and hardware.

How GLM-5.3 Flash Compares

The most useful comparison is not a single leaderboard number. It is the trade-off between capability, modalities, price, and deployment control.

ModelMain distinctionContext or pricing signalEvidence status
GLM-5.3 FlashOpen-weight, multimodal, efficiency-focused1,048,576 tokens; $0.15/$0.50 per 1M tokens on CloudflareOfficial documentation
GLM-5.3Flagship GLM-5.3 family modelFlash is positioned at roughly one-tenth the priceZ.ai relative claim
Claude Opus 4.8Closed frontier comparison modelFlash was reported at 84.3 versus 85.0 on Terminal-Bench 2.1Third-party report
Qwen3.8 FlashSpeed-focused open model comparisonCommunity testing reported approximately 50 tokens per second versus approximately 25 for GLM-5.3 Flash in proseCommunity test
DeepSeek V4 FlashLow-cost open model comparisonCommunity comparison placed it ahead on speed and below Flash on intelligenceCommunity opinion and testing

Z.ai reports 63.4 on DeepSWE v1.1 and 48.8 on AutomationBench for GLM-5.3 Flash. Those figures show the model's intended coding and agentic positioning, but independent replication and identical test configurations remain limited.

An earlier analysis examines the broader GLM-5.3 and Claude relationship: GLM-5.3 vs Claude: Open-Weight Challenger Meets Opus 4.8.

For developers evaluating a closed comparison model in a separate workflow, Claude Opus 4.8 is a relevant reference point for coding and agentic performance, not a substitute for testing GLM-5.3 Flash on the target workload.

Availability: How to Access GLM-5.3 Flash

The official access paths are Z.ai's API, GLM Coding Plan, ZCode, and Z.ai chat products. The model code documented by Z.ai is glm-5.3-flash. The API documentation supports image content blocks, streaming, tool streaming, function calling, and a 1-million-token context.

Z.ai states that GLM-5.3 Flash is fully available on its GLM Coding Plan and receives three times the quota of GLM-5.3 within that plan. Off-peak calls, including weekends, consume 50% of the standard points according to the supplied documentation.

The official weights are available in the zai-org/GLM-5.3-Flash model repository. The MIT License supports inspection, modification, and self-hosting, subject to the terms of the released weights.

Local deployment is possible, but “possible” does not mean lightweight. Community serving recipes describe a 4-bits-per-weight EXL3 build of approximately 164 GiB across two NVIDIA DGX Spark systems. That recipe reports 62.9 tokens per second for one structured stream and 26.9 tokens per second for prose under its own test conditions: the two-DGX-Spark serving recipe.

Other community reports describe Q2 or Q4 inference on a single 128 GB MacBook or DGX Spark. These are quantized configurations and should not be treated as the default requirements for every runtime.

What We Don't Know Yet

Several important engineering details remain unsettled:

  • Z.ai's direct per-token API price and complete rate-limit schedule are not confirmed in the supplied official documentation.
  • The documented 1,048,576-token context is official, but practical reports often use 256K or 262K contexts.
  • No independent evaluation yet establishes stable quality at the full 1-million-token limit.
  • Benchmark comparisons use different versions, prompts, effort settings, and serving environments.
  • The exact effect of Z.ai's August 28 configuration update on agentic workloads requires before-and-after testing.
  • Hardware requirements vary substantially by precision, quantization, KV-cache format, speculative decoding, and concurrency.
  • Data-retention policies depend on the access channel and are not a property that can be inferred from the model weights alone.

The correct engineering posture is to treat official architecture and interface specifications as confirmed, while treating throughput, benchmark rank, and cost-per-task claims as workload-specific.

Frequently Asked Questions

What is GLM-5.3 Flash?

GLM-5.3 Flash is Z.ai's open-weight, natively multimodal mixture-of-experts model for coding, agentic reasoning, and visual software workflows. It has 320B total parameters, 18B active parameters, and a 1,048,576-token context window.

Is GLM-5.3 Flash open source?

GLM-5.3 Flash is released with publicly available weights under the MIT License. The weights are available through the official zai-org/GLM-5.3-Flash model repository, while the serving stack and quantization may come from separate projects.

How much does GLM-5.3 Flash cost?

GLM-5.3 Flash costs $0.15 per million input tokens and $0.50 per million output tokens on the Cloudflare Workers AI listing. Z.ai's supplied documentation confirms low-cost positioning but does not provide a separate direct API rate in the available material.

What is the GLM-5.3 Flash context window?

GLM-5.3 Flash has a documented context window of 1,048,576 tokens, commonly described as 1 million tokens. Its maximum output length is 128,000 tokens.

Can you run GLM-5.3 Flash locally?

GLM-5.3 Flash can be run locally because its weights are public, but the full-precision model is not a lightweight desktop download. Community recipes demonstrate quantized deployments on DGX Spark systems and other 128 GB-class hardware, with results varying by quantization and runtime.

GLM-5.3 Flash vs GLM-5.3?

GLM-5.3 Flash is the lower-cost, open-weight, natively multimodal member of the family, while GLM-5.3 is positioned as the larger flagship for demanding agentic work. Flash uses 320B total parameters with 18B active and adds public weights, image input, and video input.

GLM-5.3 Flash vs Claude Opus 4.8?

GLM-5.3 Flash is positioned as a lower-cost open-weight alternative to Claude Opus 4.8 for coding and agentic workloads. Z.ai reports that Flash approaches Claude Opus 4.8 on selected evaluations, but independent, apples-to-apples replication remains limited.

What to watch next: Track Z.ai's direct API pricing and rate limits, full-context serving results near 1,048,576 tokens, and independent evaluations of coding, vision, and tool-use reliability.

Update — 2026-09-14

NVIDIA has released GLM-5.3-Flash-NVFP4, adding a vendor-provided quantized build for NVIDIA hardware alongside the community formats already available.

Community deployment tooling has also advanced. The two-DGX-Spark EXL3 project added shorter waits under long-prompt contention, configurable reasoning defaults, cleaner tool-call history, and more predictable startup behavior, according to MiaAI Lab’s September 10 update. A separate two-Spark implementation reported 1,408 tokens per second for cold prefill at 240K tokens and 69–70 tokens per second for structured decoding with speculative decoding, but these remain single-configuration community measurements, not standardized benchmarks.

Building similar long-context multimodal coding and agent workflows? On kie.ai you can try Claude Opus 5.5, Gemini 3.8 Flash, and DeepSeek-V4.1-Flash.

Maya Chen

About Maya Chen

Maya tracks AI model releases, benchmarks, and developer adoption signals across the open and closed model landscape.

View all posts by Maya Chen