Meta released Muse Spark 1.3 on September 2, 2026, rolling the model into Muse Code and the Meta Model API. The research post frames the update around longer-horizon agentic work, tighter coding workflows, and better calibration on limits — not a new open-weight release.

The launch sits on the same proprietary Spark line Brocker covered with Muse Spark 1.1 and Muse Spark 1.2 inside Muse Code; 1.3 is the next incremental drop on those surfaces.

Benchmark Muse Spark 1.3 (max) Meta Muse Spark 1.2 (xhigh) Meta GPT 5.6 Sol (max) OpenAI Opus 5 (max) Anthropic
GDPVal-AA v2 Knowledge work 1754 1615 1710 1824
JobBench Professional tool use 64.9 61.6 45.4 65.7
OSWorld 2.0 Agentic computer use 66.9 47.6 62.7 68.3
DeepSearchQA Agentic browsing 89.4 85.9 93.0 90.4
Agentic IF Index (Internal) Instruction following 57.8 46.2 60.5 59.1
AutomationBench E2E business workflows 49.4 38.2 46.7 50.3
MRCR 256K-512K Long context 98.5 66.3 91.5
MRCR 512K-1M Long context 98.1 55.5 73.8
DeepSWE v1.1 Long-horizon agentic coding 75.4 55.0 73.0 74.0
SWEAtlas CodeBase QnA Codebase understanding 59.4 46.2 53.5 52.7
Terminal-Bench 2.1 Agentic terminal coding 88.8 82.9 88.8 86.7
= not reported by vendor in this release

Figures are vendor-disclosed; higher is better unless noted. Independent replication pending.

Confirmed

  • Availability: Muse Spark 1.3 is rolling out in Muse Code and the Meta Model API as of September 2, 2026.
  • Agentic behavior: Meta says 1.3 sustains longer threads, juggles multiple workflows in one conversation, uses tools to build context across conflicting sources, asks clarifying questions on ambiguous prompts, requests help when stuck, and confirms before consequential actions.
  • Instruction following: The company reports more reliable preservation of detailed requirements across multi-step tasks versus earlier Muse Spark builds.
  • Multitasking: Meta says the model maps new prompts to the correct task more accurately when users interrupt or steer mid-thread.
  • Capability awareness: Training aimed at better judgment on what the model can and cannot do, and when to stop rather than hallucinate progress.
  • Coding efficiency (internal): Meta engineers preferred 1.3 over 1.2 as faster, with roughly 20% fewer tool calls and about 25% fewer tokens in comparisons cited in the research post.
  • Safety claims: Meta reports stronger adversarial robustness, improved prompt-injection resistance, and better calibration on irreversible actions on complex agentic tasks.
  • Reasoning modes: Previously available reasoning modes are live with 1.3; max reasoning is explicitly deferred until additional safety testing finishes.

Unknown

  • Max reasoning timing: No ship date beyond “shortly after” safety testing in the primary post.
  • Public pricing: Mark Zuckerberg posted on X that performance is “almost too cheap to meter” — the research blog does not publish API price sheets or unit economics; treat invoice impact as unverified until Meta posts numbers.
  • Vendor benchmarks: Meta’s scorecard comparing Muse Spark 1.3 (max) against GPT 5.6 Sol and Opus 5 is company-supplied; independent replication is not available in the material.
  • Demo outputs: Blog demos (CFD report, audio edit, PowerPoint, constituent summary) are labeled prototypes created by the model, not shipping products.
  • Open weights: Zuckerberg teased upcoming open-weight releases on X; the September 2 research post does not name models, dates, or licenses.

Our take

Meta’s pitch is practical agent reliability — fewer dropped constraints, fewer wasted tool calls, and explicit pauses before irreversible steps — not a leaderboard crown. That matches where production agents actually fail. The gap to watch is the split rollout: developers get 1.3 today, but the “max reasoning” tier Zuckerberg and the benchmark chart emphasize is still behind a safety gate. Until that mode ships with public pricing, “almost too cheap to meter” is marketing air, not a bill you can model. Same unit price also does not mean same invoice per job if 1.3 simply runs longer or calls more tools on hard tasks.

Series: Muse Spark 1.1 · Muse Spark 1.2 / Muse Code · Muse Spark 1.3 · Muse Spark coverage

Sources