A Nature-covered shadow evaluation and Anthropic RSI essay reveal a gap between AI engineering speed and research judgment
A recent shadow evaluation covered by Nature found an AI agent using Claude Opus 4.8 failed at open-ended research judgment despite strong engineering execution. Anthropic’s own RSI essay confirms an 8x code acceleration but admits the same judgment gap persists.