Self-Play Pretraining: Two AIs Train Each Other with Zero Human Data
A new paper from researchers at Tel Aviv University, Stanford, LAPTh and an independent researcher introduces a pretraining procedure where two transformer models, both starting from random initialization, teach each other without ever touching a single sentence of human-produced text. The result is predictable compute-scaling on natural language, images, speech, music, DNA sequences, and mathematics. If it scales, the hard scarcity in AI training may shift from data to pure compute.
Why This Problem Exists
Every major language model trained to date has relied on enormous human-curated corpora. The Common Crawl, books, code repositories, scientific papers, these are finite. Cleaning and preparing them is a separate industrial undertaking. Synthetic data generation is now common, but it typically starts with a capable pretrained model that already absorbed human text. The bootstrapping question has never been answered cleanly: can useful structure emerge with no human text in the gradient path at all?
The paper, arXiv:2609.30063, submitted September 24, 2026, by Aditya Cowsik, Kfir Dolev (Tel Aviv University), Michael Y. Li (Stanford University), G. Bruno De Luca (LAPTh, USMB), Nourya Cohen (Tel Aviv University), Noah D. Goodman (Stanford University), and Yoav Levine (Tel Aviv University), attempts to answer that question. They describe the approach as “an initial proof-of-concept towards realizing this vision,” which is an honest framing of where the research currently sits.
Their motivation follows what Richard Sutton called the Bitter Lesson: methods that rely on scalable computation tend to outrun methods that rely on hand-engineered structure. The authors extend this logic one step further: if training data itself can be generated by the training loop, compute becomes the only resource that needs to scale, rather than both compute and data acquisition.
How the Self-Play Procedure Works
The setup has two components trained from random initialization. The Learner is an autoregressive transformer that predicts byte sequences using standard next-token cross-entropy loss. The Generator is a second autoregressive transformer that produces programs for a minimal universal Turing machine. Those programs execute to produce the byte sequences the Learner trains on. The authors use a Brainf*ck-like Turing-complete language as the instruction set, which manipulates a byte-valued tape, implements loops, and reads or emits bytes. Because the language is universal, any computable data-generating process can in principle be represented as a program.
Each round of self-play proceeds through three steps. First, the Generator samples N programs. Second, each program is executed on the universal machine with a random input tape to produce an output byte sequence. Third, the Learner takes one gradient step on those sequences; the Generator takes a policy gradient step with a reward based on learning progress.
The Generator reward is the interesting part. An obvious choice would be to reward programs that produce sequences the Learner finds difficult to predict. But this has a failure mode: random byte sequences are maximally difficult to predict yet contain no learnable structure. Instead, the authors define reward as a preconditioned gradient-alignment score. Specifically, the reward for a program is the inner product of two vectors: the Learner’s gradient on that program’s output, and the Learner’s recent parameter movement over a lookback window. The inner product is computed under the diagonal AdamW preconditioner from the Learner’s optimizer state. Programs that produce gradients pointing in the direction the Learner has already been moving get high reward; programs that produce gradients in irrelevant directions get low reward. The intuition is that “learning signals arising from reusable structure accumulate, whereas idiosyncratic effects that are not learnable do not.”
The Generator is trained with a GRPO-based policy gradient estimator, with KL regularization toward a Solomonoff-inspired uniform prior weighted by program length. The program pool at each round contains fresh samples from the current Generator, local mutations of previously high-reward programs, and programs replayed from earlier rounds. Fresh samples provide global exploration, mutations refine promising regions of program space, and replay preserves structures discovered earlier.
Results: What They Measured
The authors ran a compute-optimal scaling law analysis. They trained randomly initialized transformers at various scales across multiple self-play rounds and evaluated zero-shot performance on held-out natural datasets. The datasets span natural language, images, speech, melodies, DNA sequences, and mathematical sequences. For each dataset, they constructed a compute-optimal frontier over three variables: model size, number of self-play rounds, and ensemble size.
The main finding is that zero-shot loss improves predictably with compute, following a power-law relationship. Scaling exponents are described as “comparable to those obtained by training directly on natural data” across these diverse domains. A single family of self-play models exhibits this consistency, which is the key result: a general-purpose predictive structure emerges from a compute-driven search with no domain-specific seeding.
Two secondary results strengthen the case. First, the Learner develops in-context learning on completely held-out tasks without any gradient updates after training, a capability that was not directly optimized for. Second, during training, the Generator independently discovered known mathematical sequences, including Fibonacci-like patterns, without any explicit reward for doing so. Both suggest that the training procedure, while simple in its objective, produces genuinely general-purpose predictive structure.
The scaling behavior is described as predictable specifically because it forms a clean power law rather than a one-off result at a single scale. The compute-optimal frontier shows that allocating more compute produces consistent returns across self-play rounds and model sizes.
Limitations and Open Questions
The authors are explicit about a caveat in the “zero data” framing. While no human data was used for gradient updates, some natural-data validation loss was used for hyperparameter and model-selection decisions during experimentation. The authors acknowledge this introduces limited information leakage. It is not a hidden weakness, they flag it in the paper, but it means the tabula rasa property is not absolute.
More fundamentally, the system learns generic predictive structure, not contingent world knowledge. The paper distinguishes between “contingent information, the particular facts, symbols, or modality of any one dataset” and structural regularities like copying, recursion, and hierarchical composition. Self-play over computable structure can plausibly discover the latter. Whether and how it could acquire the former remains completely unresolved.
This is an initial proof-of-concept at relatively small scale. The scaling law results are encouraging but do not yet approach frontier pretraining scale. The compute cost of the self-play loop, maintaining two models and computing preconditioned gradient alignment via forward-mode automatic differentiation, has not been analyzed against the cost of training on equivalently sized natural corpora. It may not be cheaper per bit of learned structure; it may simply scale differently.
The generator’s search space is the set of programs for a particular universal Turing machine, which is expressive but not identical to the distribution of natural data. Transfer is possible because predictive structure is separable from content, as prior work on formal-language pretraining has shown. But the gap between structural priors and world knowledge is large, and bridging it from this starting point has no demonstrated path.
What This Means for Engineering Teams
The immediate practical relevance is modest: no production system today would use this in place of a curated pretraining corpus. The paper is clear that this is proof-of-concept research. But the architectural shift it points toward matters for teams thinking about how foundation models will be built over the next five years.
If self-play pretraining scales, the data pipeline changes from an external industrial process into an optimization problem inside the training loop. Instead of a team building and maintaining a multi-terabyte cleaned corpus, the training system discovers its own curriculum. This would be a significant change in how AI model development is organized, less like traditional software data pipelines, more like a compiler that generates its own test suite. For teams currently building AI systems and pipelines, this direction suggests that compute efficiency and curriculum search become first-class engineering concerns rather than afterthoughts.
The in-context learning result also matters for applied teams. A model that develops in-context learning without gradient updates on the target task is practically useful: it suggests that structural pretraining can produce transfer-ready models before any downstream fine-tuning. If this behavior strengthens at scale, fine-tuning requirements for many tasks could decrease.
Finally, the reward design, preconditioned gradient alignment using the optimizer’s own state, is a concrete algorithmic idea that may have applications beyond pretraining. Any learning system that needs to prioritize which examples are useful at the current stage of training faces a version of this problem. The authors’ approach of using the optimizer’s momentum state as a curriculum signal is worth attention from teams building adaptive training pipelines and synthetic data generation systems.
Key Takeaways
- Two randomly initialized transformers trained via self-play, with no human data in the gradient path, produce predictable compute-scaling across natural language, images, speech, music, and DNA sequences.
- The Generator reward uses preconditioned gradient alignment with the Learner’s AdamW optimizer state to target programs at the Learner’s learning frontier, avoiding the random-noise failure mode of difficulty-based rewards.
- Scaling exponents on zero-shot natural-data loss are comparable to training directly on natural corpora, though the comparison has not been done at frontier scale.
- The Learner develops in-context learning on held-out tasks without fine-tuning; the Generator independently discovers mathematical sequences like Fibonacci patterns.
- The system acquires structural predictive priors, not factual world knowledge; the path from structural priors to general-purpose assistants remains unresolved.
- If it scales, the scarce resource in pretraining could shift from labeled corpora to compute, changing how AI training pipelines are built and staffed.
Work With Origins AI
Origins AI builds production AI systems for engineering teams. If your team is designing data pipelines, training workflows, or evaluation infrastructure for large-scale AI models, talk to our team.

