On September 8, Microsoft published benchmark results showing its Discovery Engine with CLIO (Cognitive Loop via In-Situ Optimization) outperformed other agentic harnesses on Agent's Last Exam, an evaluation of long-running, tool-using professional tasks. The system scored 61.6% in health and medicine, 75.2% in physical sciences, and 64.6% in life sciences, according to the company's Azure blog post authored by Corporate Vice President Aseem Datar.
CLIO enables independent reasoning paths to explore a problem, compare and share learning, and resolve the strongest trajectory into a single evidence-backed result. The system can determine when to keep exploring, change strategy, use a different model, or bring a domain expert into the loop. Microsoft Discovery is positioned as an enterprise platform for agentic R&D that combines hypothesis-driven experimentation with engineering rigor around reproducibility and governance.
Confirmed
- Benchmark: Agent's Last Exam, described as a demanding evaluation of long-running, tool-using professional tasks.
- Scores reported by Microsoft: 61.6% (health and medicine), 75.2% (physical sciences), 64.6% (life sciences).
- Architecture: CLIO adds an adaptive reasoning loop and diverse model ecosystem to Microsoft Discovery Engine.
- Platform availability: Microsoft Discovery is offered to R&D organizations across industries and the scientific community.
- Real-world example cited: Discovery Engine with CLIO supported work that discovered a novel organic redox flow battery.
- Potential application areas named: design simulation (silicon chips), formulation and process optimization (manufacturing, CPG), materials and molecular discovery (sustainability, drug discovery), and lab automation.
Unknown
- Independent replication of the Agent's Last Exam scores by third parties.
- Details of the benchmark's task composition, scoring rubric, and number of tasks per domain.
- Pricing, licensing, or access tiers for Microsoft Discovery with CLIO.
- Compute requirements, model lineup, or latency profiles for the CLIO reasoning loop in production.
- Whether the redox flow battery discovery has been peer-reviewed or validated experimentally beyond the blog claim.
Our take
Microsoft is framing agentic discovery as a platform play, embedding adaptive reasoning inside an enterprise R&D stack that already handles governance and data lineage. The benchmark lead is real but narrow — three domain scores from a single vendor-run evaluation — and the redox flow battery claim lacks a publication reference. The next proof points are partner integrations and evidence that domain experts actually adopt the human-in-the-loop controls rather than bypass them.