This is one outlet's own report from Time — the article as it was filed. Other outlets are covering the same event; open the full story to compare every source side by side.
See the full story · 1 sourcesInside the Race to Make AI Build Itself

This is one outlet's own report from Time — the article as it was filed. Other outlets are covering the same event; open the full story to compare every source side by side.
See the full story · 1 sources
—Getty Images
Jack Clark, Anthropic’s co-founder, left on paternity leave last November. When he returned in February, he was surprised to learn that colleagues hardly wrote code anymore. They managed five or six copies of the company’s AI, Claude, which sometimes managed several more Claudes.
To Clark, this looked like an early form of something the field has anticipated and feared for decades: recursive self-improvement, or the point at which AI begins to accelerate its own development. First, the thinking goes, models make researchers faster, but as each improvement feeds the next, the models take over more of the research cycle. Years of progress compress into months—leaving society with little time to absorb the consequences, from job disruption to engineered pathogens. Taken to its limit, AI could improve itself without humans, triggering a runaway loop long known as an "intelligence explosion," where machines rapidly advance beyond human understanding, and possibly beyond human control.
Clark believed the world needed to confront this prospect. He posted a flurry of blog posts on recursive self-improvement and flew home to England to deliver a talk on the topic in May, then led an Anthropic report in June titled “ When AI Builds Itself ,” arguing the technology is already accelerating its development. The volume of code produced per person at Anthropic has increased eight-fold, with Claude writing 80%, the report noted. “We're trying to help substantiate this concept now ... before it becomes something that is politicized or otherwise gains some valence that makes talking about it difficult,” Clark says.
It worked out as Clark feared. “Anthropic is trying to strike terror into everyone’s hearts,” wrote AI skeptic Gary Marcus, adding “all they have really shown is just faster coding.” Skeptics point out that technological progress has always compounded. Oil is used to drill oil. Why, in AI’s case, should the curve suddenly bend upward? For a company betting on continued advances, they argued, the claim is plainly self-serving.
Even Clark concedes coding volume is a crude yardstick. Claude’s code can be long-winded. But the trouble runs deeper. Neither Claude nor any large language model is written in code at all. Researchers set growth conditions—deciding the size and shape of a neural network, then pour an internet’s worth of text through it, letting it adjust itself billions of times until abilities to answer questions, write code, and hold a conversation emerge. They are cultivated, the way one grows a plant by tending the soil and the light without ever deciding where a single leaf will go.
Progress, therefore, depends on trial and error. If Claude could take over that cycle, designing, running, and analyzing experiments, would progress accelerate gradually… or suddenly explode? And if it did, could anyone pump the brakes? The uncomfortable truth is that the people building the technology are nearly as much in the dark as everyone else.
Claude began beating the benchmarks
The change Clark had walked into was not entirely unexpected. Fellow Anthropic co-founder and chief science officer Jared Kaplan had long feared that AI would eventually accelerate research, perhaps outpacing safety efforts. In early 2025, he folded a warning into the company’s Responsible Scaling Policy, its plan for managing AI’s growing dangers. Back then, Claude was no good at running experiments. But he believed that would change one day, and Anthropic would need to be ready.
To find out whether that day was coming, Anthropic built a series of tests—tasks that would take a human expert hours. Could Claude train a smaller AI model from scratch? Could it program a virtual robot dog? No single measure would settle it, but together they offered a snapshot.
In one test, Claude had to rewrite a piece of code to use GPUs—the chips for training AI—more efficiently. In spring, it momentarily got them running seven times faster, then broke the code. By summer, a newer Claude pushed the same speedup from seven times faster to 73 without introducing errors. “We started seeing these tasks fall over,” says Daniel Freeman, a member of Anthropic’s frontier red team who designed the evaluations.
As Claude began completing such tasks too reliably to reveal much, Freeman devised a test closer to the messiness of the real world. Two human teams competed in a series of challenges, involving a robotic dog and a beach ball. One team could use Claude, the other could not. After narrowly losing the first challenge, the team using Claude finished the second nearly two hours ahead. With time to spare, they trotted their robotic dog around the warehouse—until a miscalculation sent it springing toward the other team, still hunched over their laptops. An overseer caught it just in time.
The November release of Claude Opus 4.5 was a tipping point. Where previous generations tended to stall partway, the model could carry a researcher's experiment through to the end, freeing them to run more at once. "Before, I'd have eight ideas and I'd try one of them," Kaplan told TIME in February. "Now I ask Claude to just try all eight."
Ahead of a February release, Anthropic surveyed 16 of its researchers, asking whether it could replace an entry-level colleague. Five thought it might. Asked to reflect on their answer, all five walked it back. That they even entertained the idea—that they were already largely redundant—is perhaps more revealing than any benchmark.
If Anthropic’s tests fell, outside measures have similarly reached their limit. The METR graph, perhaps the best-known independent measure of AI software engineering ability, tracks the complexity of tasks models can complete based on how long they would take a human expert. But in May, Claude exceeded the benchmark’s upper limit.
This spring, Anthropic let the robot dogs out again, but this time, an improved Claude worked alone. On every task it could attempt, it was at...
AIPROPX is an independent multi-source news index — we track, compare, and connect coverage from across the web into one place you won't find anywhere else.