'Shadow Evaluations' Test Whether AI Agents Can Do Real Research — Both Attempts Would Be Rejected
Given six days and thousands of dollars in compute, frontier AI agents finished the engineering on two real research questions but couldn't answer them.