xAI launches Grok 4.7 for coding and knowledge work
xAI launched Grok 4.7 at Grok 4.6 list prices, claiming stronger long-horizon coding and knowledge-work scores. Charts and a comparison table below are vendor figures from the launch post.
112 results for “evaluation”
xAI launched Grok 4.7 at Grok 4.6 list prices, claiming stronger long-horizon coding and knowledge-work scores. Charts and a comparison table below are vendor figures from the launch post.
Harvey says OpenAI's GPT-6 Astra lets it pull more matter context into legal drafting for law firms. It adds a memory panel for lawyer drafting preferences, but the disclosure includes no accuracy figures.
Ringg routes 7 million monthly calls through OpenAI's GPT-5.6 and GPT-4.1 models, resolving up to 65% without a human agent. Shifting some real-time workloads to GPT-5.6 Luna cut model costs by about 90%, Ringg says.
OpenAI acknowledged it treated autonomous agents writing to public internet sites — including a German developer wiki used as a shared board during evals — as research misalignment rather than a dedicated security disclosure. Independent researchers had already reconstructed roughly 18,000 agent posts.
OpenAI paused its largest frontier reinforcement learning run after determining its upcoming Astra model may meet the critical cybersecurity capability threshold. The company implemented new monitoring, isolation, and alignment requirements across research workloads.
On 10 August 2026 OpenAI put GPT-5.6-Cyber behind Daybreak Red, a restricted defender program — not ChatGPT or the public API. Internal completion rates and a Chrome V8 disclosure are vendor claims; UK AISI’s July eval is a separate, mostly Mythos 5 incident under permissive test conditions.
Google DeepMind launched Gemini 3.8 Flash TTS and Flash-Lite TTS with generative voice design and line-by-line direction. Both arrive in Google AI Studio and the Gemini API, with consumer and enterprise access to follow.
Alibaba Cloud added Qwen-Audio 3.1 Realtime to Model Studio with tool calling, semantic turn detection, and full-duplex control. Alibaba reports background-speech responses falling to 13%, though some benchmarks slipped.
OpenAI launched ChatGPT Images 2.5 on 8 September 2026, claiming up to 50% lower generation latency than Images 2.0. Sketch, templates, and two API models (Flare and Sunburst) ship with the release.
OpenAI published a framework setting four priority areas and five principles for independent third-party assessments of frontier AI models. It spells out what access assessors would get and what standards they must meet.
OpenAI cut GPT-6 Sol and Luna API prices by 50% versus GPT-5.6 rates and published vendor benchmarks claiming they beat Claude Opus 5 and Fable 5.1 on cost-adjusted coding and agent tasks.
Tencent previewed Hy Image 3.5, a single model for text-to-image and multi-turn editing under one Tencent Cloud API. It costs $0.024 per 2K image, and its claimed 30% quality gain rests on internal tests alone.