Muse Code Release: Technical Deep Dive

Lukas Vogel

Lukas Vogel

Applied Research Editor

Published: September 1, 2026
Terminal coding agent workflow represented through parallel software engineering tasks and event logs

TLDRPersistent agents, a local event log, 1M-token context, pricing, and benchmark claims make Muse Code a notable coding-agent release.

Breaking Down the Muse Code Release: Persistent Agents, a 1M-Token Context, and Restart-Safe Sessions

On August 5, 2026, Muse Code entered public view as a beta terminal coding agent powered by Muse Spark 1.2. A later community report dated August 31 says the tool is now out of beta and programmable, but that status is not independently confirmed in the available documentation.

TLDR Muse Code is notable less for a single model score than for its runtime design: persistent background agents, parallel worktrees, a 1M-token context window, and a local event log intended to make long coding sessions resumable. Pricing is published, benchmark claims are circulating, and independent validation remains limited.

Key Takeaways

  • Muse Code is a terminal-based coding agent, while Muse Spark 1.2 is the coding-focused model behind it.
  • The documented runtime includes Async Background Agents, Agent Fan-Out, and a Replay-Exact Runtime.
  • The model is listed with a 1M-token context window and two API price tracks.
  • Community posts report 82.9% on Terminal-Bench 2.1 and 59% on DeepSWE 1.1, but neither result is independently validated here.
  • The strongest engineering question is durability: whether event logs and persistent agents reduce intervention without multiplying cost and risk.

What Shipped

The release introduced two related components. Muse Code is the terminal harness. Muse Spark 1.2 is the model trained for coding, debugging, codebase understanding, and end-to-end developer workflows. The official Muse Code developer documentation describes the product as a coding agent for complex workstreams.

The initial installation used one command and targeted macOS and Linux. Alexandr Wang described it as the first coding agent from Meta Superintelligence Labs, powered by Muse Spark 1.2, with access through the Meta Model API in his beta release post. Wang later described it as globally available and positioned it as one of the most affordable coding agents, but those are product claims rather than independent cost analyses.

Muse Code's published surface includes planning, code generation, validation, multimodal inputs, web search, voice mode, and multiple agents coordinating on a task. The documentation also lists a 1M-token context window for both muse-spark-1.2 and muse-spark-1.2-contributor.

The subscription page lists three plans:

PlanMonthly pricePublished allowance or positioning
Everyday Usage$5 per month10–50 requests every 5 hours, including image and video uploads
High Usage$15 per month3× more usage than Everyday Usage
Power Usage$50 per month10× more usage than Everyday Usage, plus early feature access

These figures describe the product plans. They should not be confused with token pricing for direct model calls.

The model pricing table is more revealing:

Model routeContextInputCached inputOutput
muse-spark-1.21M tokens$1.25 per million tokens$0.15 per million tokens$4.25 per million tokens
muse-spark-1.2-contributor1M tokens$0.10 per million tokens$0.002 per million tokens$0.20 per million tokens

The contributor route is marked as being used to improve products. The standard route is marked as not being used for that purpose. That distinction belongs in any production review because the nominal input price differs by 12.5×, while the output price differs by 21.25×.

Why This Release Matters

The coding-agent market has moved from line completion toward delegated work. A capable terminal agent must inspect files, form a plan, edit multiple components, run tools, interpret failures, and preserve enough context to continue. Muse Code's design directly targets that longer sequence.

The important architectural choice is persistence. The official description says background agents remain active throughout a session rather than being spawned for individual subtasks. That could reduce repeated repository investigation. It could also create additional hidden work, token consumption, and coordination complexity. The available evidence does not yet show the balance.

The second choice is operational traceability. Every model call, tool run, approval, and edit is appended to a local event log before execution. That gives engineers a chronological record of what the agent attempted. It also creates a concrete place to inspect failures, approvals, and unintended changes.

The Local Event Log is the central audit mechanism that records model calls, tool actions, approvals, and edits before they execute. That sentence is independently quotable because it describes a specific runtime behavior, not a vague promise about transparency.

This approach matters for long-running tasks. Meta's research announcement describes GPU-kernel optimization involving more than 1,000 tool calls and runs lasting up to 24 hours. Those results are vendor-reported, and the underlying evaluation report is not included in this signal set. Still, the example clarifies the target workload: iterative engineering rather than one-shot code generation.

How Muse Code Works

Muse Code's feature vocabulary describes a particular agent architecture. Several terms are useful because each maps to an observable implementation choice.

Async Background Agents are persistent supporting processes that gather information and perform follow-up work while the primary agent advances the user's task. The design reportedly aims to reduce redundant investigation and the need for constant steering.

Agent Fan-Out is the practice of sending sufficiently large tasks to multiple subagents working in parallel inside isolated worktrees. A community post attributed to Mark Zuckerberg says Muse Code built six game features simultaneously without collisions. That example is not an independent systems evaluation, but isolated worktrees provide a testable mechanism for reducing mid-task interference.

Replay-Exact Runtime describes the event-log design in which each action is appended before execution. If the implementation behaves as documented, replaying the log should reproduce the session's state transitions more reliably than reconstructing a failed conversation from prompts alone.

Restart-Safe Sessions are sessions that can resume after a crash from the recorded log rather than restarting from a blank context. This feature may matter more than raw generation speed for tasks that run for hours.

Approval-Gated Planning comes from the bundled /plan skill. The command turns a task into a plan that requires approval before execution. The related /grill skill stress-tests the plan, while /goal directs the agent toward successful completion.

Long-Horizon Coding is the training focus described for Muse Spark 1.2. The documented training examples include whole-repository generation, large end-to-end projects, and auto-research. Planning, goal conditioning, and context compaction are presented as mechanisms for maintaining direction.

Co-Training With Muse Code refers to training the model alongside the agent harness. The release materials describe rejection-sampled harness trajectories, recipe optimization for goals and subagents, and direct integration of the toolset. This could improve compatibility between the model and its runtime, although no controlled ablation is available in the signal set.

The architecture therefore creates a different evaluation target. Engineers should measure the complete loop, including planning, tool use, retries, validation, recovery, and human intervention. A model score alone cannot capture those costs.

Muse Code vs Claude Code: What the Signal Says

The community has repeatedly framed Muse Code as a competitor to Claude Code and OpenAI Codex. One post calls it a Claude Code and Codex competitor, but provides no controlled comparison. The more substantive numerical comparison concerns Claude Code running Opus 5.

DimensionMuse CodeClaude Code signal
RuntimeTerminal agent with persistent agents, parallel worktrees, validation, and a local event logThe signal set names Claude Code as a comparator, but provides no detailed runtime specification
Terminal-Bench 2.182.9% in a community report86.7% for Claude Code running Opus 5 in the same report
Context window1M tokens in the published Muse documentationUnverified in this signal set
Published model pricing$1.25 per million input tokens and $4.25 per million output tokensUnverified in this signal set

The Terminal-Bench figures come from a Reddit post containing the comparison. The post is useful because it supplies exact numbers, but it does not establish that both systems used identical prompts, tools, time limits, retry policies, or model settings.

At 82.9% on Terminal-Bench 2.1 versus 86.7% for Claude Code running Opus 5, Muse Code appears competitive on this reported test, but the result is not independently verified. This is a benchmark anchor, not a market-wide verdict.

The comparison also shows why product architecture deserves separate treatment. Claude Code's signal in this bundle is primarily a benchmark baseline. Muse Code's differentiator is its documented runtime: persistent agents, event logging, isolated worktrees, and recovery. Neither side has enough public evidence here for a complete total-cost or reliability comparison.

What the Benchmark Claims Actually Show

Two benchmark claims are circulating. A TestingCatalog post reports that Muse Code scores 59% on DeepSWE 1.1 and says the result exceeds Grok Build 4.5 and Gemini 3.6 Flash.

BREAKING 🔥: Meta AI released Muse Code, a new terminal coding agent powered by the newest Muse Spar

Source: @testingcatalog

A separate third-party newsletter summary lists Muse Spark 1.2 at 59.3% on DeepSWE 1.1, compared with 65.0% for Opus 5 and 64.8% for GPT-5.6 Terra. Because the underlying report and methodology are not present in the accessible signal set, those extra decimal figures should remain unverified.

The numbers are not necessarily contradictory. A displayed 59% may be a rounded version of 59.3%. The key issue is not rounding. It is provenance. The available material does not specify the task mix, evaluation harness, sampling settings, tool permissions, timeouts, retry policy, or contamination controls.

The same caution applies to Terminal-Bench 2.1. The reported 82.9% and 86.7% figures offer a useful lead for replication. They do not show how the systems behave on a team's own repository, with its undocumented conventions and deployment constraints.

No independent benchmark result in this signal set establishes that Muse Code is more reliable, less expensive, or safer than Claude Code, Codex, or other coding agents. The evidence supports a narrower statement: Muse Code has entered the discussion with measurable reported results and a runtime built for extended tasks.

What We Know vs. What We Don't

The boundary between documentation and inference is important. The following known items are published product details or directly reported release facts. They should not be read as independent validation.

What we know:

  • Muse Code first appeared on August 5, 2026, as a beta terminal coding agent powered by Muse Spark 1.2.
  • Published documentation lists macOS and Linux support, plus one-command installation.
  • Muse Code is described as planning changes, writing code, validating results, and coordinating persistent subagents across large repositories.
  • The runtime appends model calls, tool runs, approvals, and edits to a local event log before execution.
  • The bundled /plan, /grill, and /goal skills provide approval-gated planning, plan stress-testing, and objective-driven execution.
  • Muse Spark 1.2 is listed with a 1M-token context window, three subscription tiers, and two model pricing routes.

What we don't know:

  • The signal set contains no independent verification of the 82.9% Terminal-Bench 2.1 or 59% DeepSWE 1.1 claims.
  • The complete evaluation methodology, including task selection, harness settings, retries, and contamination controls, remains unspecified here.
  • Average latency, tool-call volume, cache behavior, and total cost for ordinary persistent-agent sessions are not established.
  • The local event log is described technically, but its complete security, retention, privacy, and repository-data policy is not available in these signals.
  • A community report says Muse Code is out of beta and programmable on August 31, but no official SDK changelog or stability guarantee appears in the bundle.
  • The available evidence does not provide a verified language matrix or results across different monorepos, build systems, and dependency graphs.

This distinction gives builders a practical reading of the release. The architecture is known. The production envelope is not.

What Builders Should Do Today

First, separate the harness from the model in internal tests. Run Muse Code on representative repository tasks, then record whether failures come from code generation, planning, tool selection, context loss, or coordination. The 1M-token context window may help with large repositories, but context capacity is not the same as useful repository understanding.

Second, inspect the event log as a security artifact. Confirm which model calls, tool runs, approvals, and edits are recorded. Check whether secrets, environment variables, file contents, and command output can enter that log. A restart-safe session is valuable only if its recovery state is safe to retain.

Third, measure the economics of persistence. Compare the standard route at $1.25 per million input tokens and $4.25 per million output tokens with the contributor route at $0.10 and $0.20. Include cached input at $0.15 or $0.002 per million tokens. The contributor route's data-use condition should be reviewed separately from the price advantage.

Fourth, test parallelism with isolated worktrees. Give separate workers tasks that touch related files, then inspect merge quality, duplicated investigation, and reviewer workload. The official documentation advertises a four-profile orchestration cookbook, while community material describes six simultaneous game features. Those examples suggest a test shape, not a guaranteed production result.

For a comparable coding-agent baseline, teams can run the same repository tasks through OpenAI Codex and compare completion quality, intervention count, recovery behavior, and total token usage. The comparison should keep prompts, tool permissions, timeout policies, and review criteria as consistent as possible.

The Week Ahead

The next useful signal is not another isolated leaderboard screenshot. It is a reproducible evaluation report showing how Muse Spark 1.2 performs with and without the Muse Code harness. An ablation between persistent background agents, parallel worktrees, and ordinary single-agent execution would clarify where the gains originate.

The second signal is documentation around programmability. The August 31 community report says the SDK can expose sessions, tools, permissions, persistent state, subagent coordination, inter-session messaging, and rewind behavior. If those details receive official documentation, engineers can assess whether Muse Code is a terminal product or a broader agent runtime.

The third signal is operational evidence. Watch for an official model card or evaluation report, run the same long-horizon task with a controlled tool budget, and pin the exact muse-spark-1.2 route before comparing results. Those three checks will reveal more than another uncontextualized score.

Building similar persistent-agent coding workflows? On kie.ai you can try Claude Opus 5.5, Claude Sonnet 5.5, and DeepSeek-V4.1-Flash.

#muse code#muse code release#muse code deep dive#muse spark 1.2#coding agent benchmark#terminal coding agent#persistent coding agents
Lukas Vogel

About Lukas Vogel

Lukas reads the papers and model cards so you do not have to, focusing on reproducible claims.

View all posts by Lukas Vogel