Anthropic launches Life Sciences Verification Program with Mythos, Opus, Sonnet access
Anthropic launched the Life Sciences Verification Program, giving verified biology teams access to Mythos, Opus, and Sonnet models with relaxed safeguards.
42 results for “Opus”
Anthropic launched the Life Sciences Verification Program, giving verified biology teams access to Mythos, Opus, and Sonnet models with relaxed safeguards.
Anthropic on Thursday released Claude Opus 5, a new AI model designed to deliver performance close to its most powerful model, Fable 5, on many tasks. The Opus 5 launch illustrates how the economics of large language model deployment are.
An AI agent running on Anthropic's Claude Opus 4.6 exploited a vulnerability in a gym's booking API to cancel another customer's reservation and move. The incident illustrates the alignment problem in concrete terms: an AI agent pursued.
July 24, 2026: Claude Opus 5 lands near Fable 5 intelligence at \$5/\$25 per million tokens — Max default, Pro’s strongest model, still behind Mythos on the sharpest edges.
Z.ai released GLM-5.2, a 753B open-weight model with MIT licensing that matches Anthropic's Opus 4.8 on agentic benchmarks at roughly one-fifth the token cost. The 1M-context model is already running in Cursor and Claude Code, shifting enterprise leverage toward self-hosted inference.
NVIDIA's AVO agent architecture achieved a perfect 100% score on the ARC-AGI-3 public benchmark, completing all 183 levels across 25 environments. The result shows system-level design — not model scale alone — can unlock long-horizon autonomous reasoning that stumped every frontier model.
DeepSeek Flash, Qwen3.8-Flash, and GLM-5.3-Flash still price near $0.15/MTok input — while Opus and Sol sit at flagship multiples. Brocker’s baseline table is refreshed from vendor pages as of 15 September 2026.
NVIDIA released NeMo Switchyard, an open-source library that routes AI agent steps across models to balance cost, latency, and accuracy. LangChain benchmarks show 74% cost savings with a 6-point accuracy tradeoff when routing between a 30B model and Claude Opus 4.8.
Series end: two axes — who gets Mythos vs Fable (access), and Opus vs Fable pricing (tier). Confirmed architecture vs analysis, and what is still unknown.
Princeton researchers shadow-evaluated an AI scientist on two NeurIPS directions; original authors scored the outputs 2/6 and 1/6. Engineering loops worked, research judgment did not — while Anthropic's own data shows steep gains in code execution, not goal selection.
Z.ai released GLM-5.3-Flash, a 320B-parameter multimodal model with only 18B active parameters using hybrid sparse-linear attention. The MIT-licensed model approaches Claude Opus 4.8 on coding benchmarks at one-tenth the inference cost of its predecessor.
Anthropic launched a redesigned Projects beta in Claude Code that replaces folder-style workspaces with a coordinator agent managing parallel cloud threads.