A Nature-covered shadow evaluation and Anthropic RSI essay reveal a gap between AI engineering speed and research judgment
Princeton researchers shadow-evaluated an AI scientist on two NeurIPS directions; original authors scored the outputs 2/6 and 1/6. Engineering loops worked, research judgment did not — while Anthropic's own data shows steep gains in code execution, not goal selection.