Skip to content
Main Site News Console

Evolution of Agent Toolchains

· Latent Space Translated
播客深度访谈

Around Christmas 2025, AI engineers noticed that agents had changed. They had started to actually work! It is difficult to pinpoint exactly why. Perhaps we finally had some time off during the holidays to experiment with the latest models and the latest agents. Perhaps the models had crossed some capability threshold. Perhaps the scaffolding around the models had finally matured.

The argument I want to make in this article is that the latter two factors worked together: the models and their runtime frameworks improved in tandem, and their improvement curves converged at just the right moment. This dynamic also helps explain what will happen next: models will continue absorbing their runtime frameworks into their weights, while engineers will continue removing the parts that have already been absorbed. Eventually, what remains will be a runtime framework designed to serve human attention, rather than one designed to serve the model.

Lukasz Kaiser, one of the inventors of the Transformer, said on the June episode of Unsupervised Learning:

“Something changed last winter, around Christmas—it’s hard to say exactly what. I mean, the runtime framework changed, post-training changed in some ways, and then new pretrained models appeared… But it felt like a huge leap, and it’s difficult to pinpoint exactly what caused it.”

The answer to “What happened?” lies not only in the model weights, but also in the entire system that has evolved around those weights.

The answer lies in the agent harness.

Think back to November 2022, when ChatGPT was the most advanced AI tool available. Its only capabilities were next-token prediction and some reinforcement learning from human feedback (RLHF) that enabled it to act like a useful assistant. It had no tools, no search, and no reasoning capabilities.

The original ChatGPT was confined entirely to its training data and the prompts you sent it—no more, no less. It was like a brain in a vat.

Agent harnesses enable large language models (LLMs) to break free from these constraints and interact with the real digital information space.

What Exactly Is a Harness?

An agent harness is everything that enables an agent to operate beyond the model weights themselves. It includes the environment, tools, context, and guardrails surrounding the model. Without a harness, a model is a brain in a vat. It can take cognitive actions, but it needs a harness to turn those decisions into actions in the real digital space.

A harness is like giving the model’s mind a body. With a harness, a model can perceive information (context), take action (tools), persist information (memory and compression), and constrain its own boundaries (permissions and guardrails).