Skip to content
Agents · Aug 22, 2026

Why simulation is replacing human roles across the AI development pipeline

From reward signals to physical-world experiments, synthetic components are becoming load-bearing in AI systems, lowering costs and accelerating timelines while trading off some quality.

Trust78
HypeLow hype

1 source · cross-referenced

ShareXLinkedInEmail
TL;DR
  • Model-generated reward signals replaced human judgments as early as 2022, becoming standard by 2025.
  • Training data, teachers, curricula, researchers, environments, and even human subjects are now routinely synthesized in frontier labs.
  • Simulation-driven approaches are expanding into physical-world experiments, particularly in bio research.

In 2022, the first component to flip from human-made to model-made was the reward signal. InstructGPT introduced a reward model trained on human preferences, allowing policies to optimize against a synthetic judge rather than humans. Constitutional AI and later work showed AI-generated feedback could match human feedback at a fraction of the cost, and by 2025, LLM-as-judge had become the default evaluation methodology in benchmarks such as MT-Bench and AlpacaEval.

By 2023, synthetic training data became load-bearing. Microsoft’s Phi series demonstrated that small models trained on LLM-synthesized, textbook-quality data could outperform expectations for their size, and Apple’s WRAP generalized this approach by rephrasing the entire web with an LLM to improve pretraining efficiency. By 2025, reasoning-trace corpora generated by strong reasoners were standard ingredients in pretraining and mid-training pipelines.

Also in 2023, the role of the teacher flipped to synthetic. Stanford’s Alpaca showed that a $600 fine-tune on GPT-generated instructions could replicate much of a frontier model’s behavior, and subsequent work like Vicuna, Orca, and DeepSeek-R1’s distilled models made model-as-teacher the default assumption for many small model releases.

In 2024, models began generating their own curricula. Meta’s Self-Rewarding Language Models and SPIN demonstrated that models could create their own tasks, judge their own outputs, and improve beyond the ceiling of human preference data, turning curriculum design—historically a human-driven process—into a self-improving loop.

In 2026, the role of the researcher flipped. DeepMind’s AlphaEvolve evolved new algorithms in 2025, and Sakana’s AI Scientist automated the full paper-writing pipeline. Karpathy’s autoresearch in March 2026 introduced a minimal ratchet loop where a coding agent iteratively modified a training setup, ran five-minute experiments, and retained changes only if validation loss improved, stacking 700 experiments into 20 kept improvements and cutting time-to-GPT-2 from 2.02 to 1.80 hours.

Also in 2026, the environment itself became synthetic. Z.ai built pipelines that synthesize end-to-end research environments by mining real work patterns, converting them into long-horizon tasks with hidden state, and validating them with judge agents and synthesized verifiers. The GLM-5.3 release described an environment, judging, and verification stack that is synthetic all the way down. Ornith-1.5 claimed end-to-end self-improvement by having the model propose its own tasks and generate its own RL rollouts.

The human subject layer is now being simulated. Simile post-trains frontier models on interviews, transaction data, and registered RCTs to recover human-like biases and causal texture, aiming to simulate real people for research and product development. Early scaling laws suggest digital twins can reproduce survey and behavioral responses with 85% accuracy compared to the humans themselves two weeks later. SimGym at Shopify and Tencent’s billion-persona approach extend this trend into commercial simulation workloads.

Simulation is also moving into the physical world. Poolside argues that durable AI value accrues to those who own the experimental loop, especially in domains like bio research where real-world experiments remain essential. CZ Biohub is building a virtual cell atlas as a precursor to a virtual immune system, citing cost and speed advantages roughly 1000x over in vivo methods.

Sources
  1. 01Latent Space — swyx[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over
Also on Agents

Stories may contain errors. Dispatch is assembled with AI assistance and curated by human editors; despite the trust-score filter, mistakes happen. We correct publicly — every article links to its revision history. Nothing here is financial, legal, or medical advice. Verify before relying on any claim.

© 2026 Dispatch. No ads. No sponsorships. No paid placement. Reader-supported via Ko-fi.

Built by a person who cares about honest AI news.