tezvyn:

AI May Automate AI R&D by EOY 2028

AI-drafted, machine-checkedSource: Import AI (Jack Clark)intermediate

Claude Mythos Preview now solves 93.9% of real-world GitHub issues on SWE-Bench, a massive leap from Claude 2's 2% in late 2023. This near-saturation of coding benchmarks is a key indicator that AI can automate its own engineering. Based on this trend, Anthropic's Jack Clark predicts a 60%+ chance of no-human-involved AI R&D by EOY 2028. This shifts the focus from AI-assisted coding to fully automated AI development.

### Why it matters Anthropic's Jack Clark predicts a greater than 60% chance that AI systems will be able to autonomously build their own successors by the end of 2028. This concept, known as recursive self-improvement, would mark a fundamental shift in technological development. The primary driver is the rapid maturation of AI's coding and task-chaining abilities, suggesting the core engineering work of AI R&D is on the verge of being fully automated.

For engineers, this moves beyond using AI as a copilot for discrete tasks. It points to a near future where AI agents could manage the entire development lifecycle, from identifying research paths to coding, testing, and deploying new models. This could lead to an exponential acceleration in AI capabilities, crossing a "Rubicon into a nearly-impossible-to-forecast future."

### What changed * **Coding Proficiency:** Claude Mythos Preview achieved a 93.9% success rate on SWE-Bench, a benchmark for solving real-world GitHub issues. This effectively saturates the test, compared to Claude 2's ~2% score in late 2023. * **Automation of Engineering:** AI systems are no longer just writing code snippets; they are becoming capable of handling complex, multi-step engineering workflows, including writing tests and checking code without human oversight. * **Expert Forecast:** Based on these public trends, Jack Clark forecasts a 60%+ probability of a no-human-involved AI R&D cycle by EOY 2028. * **Short-Term Outlook:** A proof-of-concept, where a non-frontier model end-to-end trains its successor, is considered plausible within the next one to two years.

### What to watch * **Proof-of-Concept Systems:** Watch for initial demonstrations of models training their own successors, even at a smaller scale. These will likely be the first concrete examples of this paradigm. * **Tooling and Agentic Workflows:** Monitor how frontier labs and MLOps platforms shift from providing AI-assisted tools to offering fully autonomous AI agents for R&D tasks. * **Benchmark Evolution:** Keep an eye on benchmarks like METR's time horizons plot, which measure an AI's ability to complete complex tasks that require hours of human effort.

Read the original → importai.substack.com

Get five bites like this every day.

Tezvyn delivers a daily feed of 60-second tech bites with quizzes to lock in what you learn.