Anthropic Automates AI Safety Research with Claude
Anthropic's automated AI agents, using Claude, achieved a 0.97 Performance Gap Recovered (PGR) score on a weak-to-strong supervision task, crushing the 0.23 score achieved by human researchers. This is one of the first concrete examples of automating open-ended AI research, where agents autonomously proposed, tested, and iterated on ideas. Engineers should anticipate R&D cycles accelerating as AI agents begin to tackle complex research problems.
### Why it matters Automating AI research itself has been a long-sought goal, and this Anthropic experiment provides concrete evidence that it's becoming practical. An AI agent, using Claude, was tasked with an open research problem and not only succeeded but significantly outperformed human experts. This suggests the pace of AI innovation could accelerate dramatically as the process of discovery is automated. The experiment cost just $18,000 ($22 per AAR-hour), demonstrating the economic viability of this approach for tackling complex problems.
### What changed Anthropic built "autonomous AI agents that propose ideas, run experiments, and iterate on an open research problem." The agents were tasked with improving weak-to-strong generalization, where a weaker model trains a stronger one.
* **Task**: Automate research on weak-to-strong supervision. * **Models Used**: Qwen 3-4B-Base (strong model) and Qwen 1.5-0.5B-Chat (weak teacher). * **Human Baseline**: Two researchers worked for 7 days, achieving a Performance Gap Recovered (PGR) score of 0.23. * **AI Result**: The automated agents worked for 5 days (800 cumulative hours) and achieved a final PGR of 0.97, nearly closing the entire performance gap. * **Generalization**: The AI's best method also generalized well to new datasets, achieving PGRs of 0.94 on math and 0.47 on coding tasks.
### What to watch * The application of this automated research paradigm to other, more complex AI problems beyond weak-to-strong supervision. * The development of platforms that allow engineers to deploy their own Autonomous AI Researchers (AARs) for custom R&D tasks, fundamentally changing how products are built.
Read the original → importai.substack.com
Get five bites like this every day.
Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.