xAI launches Grok 4.7 for coding and knowledge work
xAI launched Grok 4.7 at Grok 4.6 list prices, claiming stronger long-horizon coding and knowledge-work scores. Charts and a comparison table below are vendor figures from the launch post.
3 results for “LLM benchmarks”
xAI launched Grok 4.7 at Grok 4.6 list prices, claiming stronger long-horizon coding and knowledge-work scores. Charts and a comparison table below are vendor figures from the launch post.
Cursor has been acquired by SpaceX, completing an April partnership and gaining access to the largest GPU fleet for model training. Grok 4.6 launches as the first joint model, matching GPT-5.6 Sol on key benchmarks and available now in Cursor.
Amazon Science released SOP-Bench, an open benchmark with 2,000+ tasks from real business SOPs across twelve domains. The framework pairs authentic procedures with working tools and ground-truth answers so teams can validate agents before production deployment.