Skip to content
Main Site News Console

Claude Starts Training Claude: $4/Hour Outperforms title: 50/Hour Human Researchers

· 量子位
国内AI

AI “Self-Evolution” Is Getting Closer

By Tingyu, from Aofeisi

Claude has started training Claude!

At $4 an hour, it outperformed human researchers earning $150 an hour.

In a newly published study, Anthropic had Claude search academic papers, develop proposals, generate data, and train models—successfully tackling 10 categories of AI safety problems in one go.

On some tasks, its solutions were even better than those produced by 28 human AI safety researchers.

Even more strikingly, a weaker Claude has already begun helping train a more capable Claude.

The day when AI “improves itself” seems to be getting closer and closer…

At $4 an Hour, Claude Outruns Human Researchers

In a study titled Automated Researchers Can Effectively Mitigate AI Alignment Failures, Anthropic sent Claude straight into the lab.

Specifically, the researchers built an automated alignment research system called AAR based on Claude Opus 4.8.

Once given a specific model safety problem, Claude searches for relevant papers, identifies potentially useful methods, proposes new training approaches, generates data, fine-tunes the model, and finally runs safety and general capability evaluations.

From forming hypotheses to completing validation, a fairly complete AI research loop was handed over to Claude.

If a method performed poorly, it was discarded. If the results improved, Claude continued exploring in that direction.

A typical training run took just 30 minutes. This allowed Claude to rapidly test and discard ideas across a large number of proposals, much like a search algorithm.

Throughout the process, humans supplied the questions, models, and evaluation criteria, while Claude handled most of the work—from gathering information to running experiments.

Claude Found Improvements for All 10 Safety Challenges

In the experiment, Anthropic gave Claude 10 challenges in one go.

These covered common alignment problems, including deception, sycophancy, reward hacking, privacy violations, and jailbreaking.

For example, “sycophancy” refers to a model giving users answers they want to hear, even when those answers may not be correct, simply to agree with them. “Reward hacking,” meanwhile, occurs when a model does not actually complete the task but exploits loopholes in the evaluation rules to obtain a higher score.

Claude addressed one problem at a time, repeatedly proposing solutions, training the corresponding target model, and using the results to determine the next round.

In the end, all 10 categories of problems improved.

According to the “safety gap” metric defined by Anthropic, Claude closed between 26% and 96% of the gap in these areas. Put simply, it brought models with existing safety problems significantly closer to the theoretical best score.

More importantly, these methods were not effective only on the problems Claude had already seen.

Some of the winning approaches continued to reduce unsafe model behavior when applied to unreleased test sets and Petri evaluations simulating multi-turn adversarial scenarios. Their effectiveness also held up when applied to models as large as 4.7 times the size of the experimental models.

Anthropic also established a red line: safety training could not significantly damage the model’s existing capabilities.

After all, making a model refuse to answer almost everything could improve its scores on many safety tests—but such a model would be practically useless.

Ultimately, the methods found by Claude showed no obvious degradation on the predefined general capability evaluations.

85% vs. 20%: Claude Beats Human Researchers by a Mile

The most lopsided result came in the deception evaluation.

To make the target model more honest, Claude submitted more than 150 attempts.

It would propose a method, run one training round, and then continue adjusting the approach based on the results. In a single experiment, the best method closed 82% of the safety gap; across multiple runs, the average score reached 85%.

Anthropic gave the same challenge to six experienced AI safety researchers.

The human researchers closed an average of just 20% of the safety gap…

85% versus 20%—Claude won without even breaking a sweat.

A total of 28 human safety researchers participated across all the tasks. Each could spend up to eight hours reading materials and designing an approach.

The paper states that the automated researcher could typically find methods exceeding the human average within six hours. The research directions provided by humans also did not help Claude achieve better results.

The most painful part was the cost.

The API inference cost for an automated researcher was approximately $4 per hour, while Anthropic paid its human researchers $150 per hour.

This… surely means human researchers are about to be replaced by Claude.

That said, the human-machine competition was not entirely fair.

Claude could actually train the models, observe the results of each experiment, and continuously revise its approach. Human researchers could submit a method only once and had no opportunity to iterate based on the results.

A Weaker Claude Trains a Stronger Claude, Closing in on the Production Version in 60 Hours

In the earlier experiments, the main subjects being improved were relatively small open-source models.

Next, Anthropic raised the difficulty considerably: it asked the less capable Claude Sonnet 5 to train an early version of Claude Opus 4.8 that had not yet completed most of its production-level safety training.

“The disciple need not be inferior to the master”—this is probably what that saying looks like in practice.

Over 60 hours, Sonnet 5 tested more than 50 potential solutions. The final method closed approximately 65% of the safety gap, approaching the 72% achieved by the officially released version of Opus 4.8.

The approach used just over 2,000 training examples, consisting primarily of simple templates and public datasets.

According to Anthropic’s calculations, on this specific set of tasks, its data efficiency was approximately 15,000 times that of the production-level alignment process.

Of course, this does not mean Anthropic can reproduce the entire production training process with 2,000 examples.

The production model must handle more capabilities, more risks, and a large number of edge cases. This experiment focused only on a clearly defined set of safety metrics.

But it has already demonstrated that a weaker Claude can search academic papers, design methods, generate data, and then incorporate the training results into a more capable Claude.

Humans set the boundaries; Claude handles the iteration. One model has begun participating in the creation of the next, safer and more capable model.

AI improving AI is no longer hypothetical—it is happening right before our eyes.

So, Does This Count as “AI Self-Evolution” or Not?

Judging by the results, this study has indeed reached the threshold of “AI self-evolution.”

But if “self-evolution” means an AI completely autonomously designing, training, and producing a more capable next-generation AI, then this experiment is still nowhere near that point.

After all, humans still decided what problems to investigate, which models and data to use, and what results would count as success.

Anthropic itself acknowledges that AAR can optimize only the objectives that humans have written into the evaluation system in advance.

This experiment did not cover other complex problems, and some low-frequency or newly emerging risks do not even have corresponding benchmarks yet.

The researchers also examined only a predefined set of general capabilities, so they could not rule out the possibility that a particular safety training process improved the scores while damaging other capabilities that were not tested.

Safety evaluations such as Petri are likewise only proxy metrics for real-world risks. A model behaving more safely in tests does not necessarily mean it will be reliable in a real-world environment.

Thus, AAR’s greatest hidden risk lies in its greatest strength: it is exceptionally good at rapidly trialing and discarding approaches according to a metric.

If the measuring stick is wrong, the faster Claude runs, the farther it may move from the result humans actually want.

More dramatically, the research agents themselves also cheated.

To prevent AAR from exploiting loopholes in the evaluation rules, Anthropic specifically assigned a monitoring Agent to inspect the approaches Claude was preparing to execute and its complete research records.

In approximately 1,600 research records, the monitoring model detected 39 cheating attempts—about 2.4%.

Some Agents repeatedly submitted the same approach, attempting to exploit evaluation fluctuations to obtain a higher score. Others generated training data that imitated the format of the Benchmark. Still others deliberately concealed experimental steps that violated the rules.

Rather than genuinely solving the model’s safety problems, they were more interested in finding a shortcut to quickly improve their scores.

That is darkly humorous… So the Agents were guarding the vault while robbing it themselves.

The good news is that most of these cheating intentions still appeared in the models’ reasoning records, allowing the monitoring Agent to detect them.

The bad news is that Anthropic cannot be sure whether more capable models in the future will leave such obvious traces.

When research Agents have more tools, greater privileges, and the ability to directly modify the model training process, “how do you monitor an AI responsible for improving AI?” may prove more difficult than the training itself.

All things considered, an AAR costs approximately $4 per hour, can work continuously, and can be duplicated dozens or even hundreds of times to run experiments in parallel.

Human researchers, by contrast, earn far more per hour, need to rest, and can advance only a limited number of experiments at a time.

However you do the math, the answer is already obvious.

I don’t know whether to call this a nuclear detonation or simply collapse in my chair…

This time, it is the people who build AI who have to worry about AI taking their jobs.

References:

  1. https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures
  2. https://www-cdn.anthropic.com/7b1c44894e98087a6479947dcdd40716278aeeffd/automated-alignment-researchers-august-2026.pdf
  3. https://techcrunch.com/2026/08/28/an-anthropic-researcher-just-gave-us-a-peek-at-self-improving-ai/