Skip to content
Main Site News Console

[AINews] 10% Lower Performance, 100× Lower Cost, 10,000× Faster: Why Simulation Is Taking Over

· Latent Space Translated
播客深度访谈

By today’s AI industry standards, this is a remarkably quiet Friday, so it’s time to take a step back and reflect on what exactly has happened. If you’ve read our “2025 Reading List”, followed our coverage of Z.ai GLM, understood Poolside’s pivot, kept up with our AI for Science theme, and listened to today’s Simile podcast, then you’re not only part of Latent Space’s largest readership group—you may also have arrived at a mental model that looks something like this:

Image

Since 2022, one component of the pipeline for generating machine intelligence has shifted each year from being human-made to being model-made. This shift has been neither gradual nor uniform—each transition had a “Patient Zero”: a paper or product in which the synthetic version first became load-bearing at a frontier lab. From that point on, the future had already arrived; it simply had yet to be productionized.

If you squint a little, you’ll see that what we used to call “synthetic data,” “synthetic rubrics,” “AI researchers,” and “end-to-end reinforcement learning environments” are all, at their core, increasingly ambitious simulations of humans—10% worse in quality, but 100 times cheaper and 10,000 times faster.

Stage 1: Reward Signals (2022)

The first thing to become synthetic was, surprisingly, the evaluator. InstructGPT established what has since become the canonical approach: collect human preferences once, train a reward model, and then optimize the policy against the model rather than against humans. From the policy’s perspective, the entity handing out approval was already an LLM.

Constitutional AI took this a step further by having AI self-criticize according to a set of principles (RLAIF); later, Lee et al. showed that AI feedback could match human feedback at a fraction of the cost. By the time LLM-as-judge became the default evaluation method (MT-Bench, AlpacaEval), the entire approval mechanism—reward, criticism, and evaluation—had become models judging models.

Stage 2: Training Data (2023)

Microsoft’s Phi series put the argument right in its title: “Textbooks Are All You Need”. A small model trained on LLM-synthesized, textbook-quality data performed far beyond what its parameter count would suggest, and phi-1.5 confirmed that this was no fluke.

Apple’s WRAP took the approach even further: don’t just generate data—have an LLM rewrite the entire internet—increasing pretraining efficiency by roughly 3×. From there, the pipeline began to industrialize—NVIDIA released Nemotron-4 340B with a permissively licensed synthetic data generation pipeline as one of its primary selling points; by 2025, מחדש