A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Beyond Model Collapse: Scaling Up with Synthesized Data Requires Reinforcement
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
Grandmaster level in StarCraft II using multi-agent reinforcement learning
β scaling-hypothesis#blessings-of-scale
Diversifying AI: Towards Creative Chess with AlphaZero (AZdb)
Best Practices and Lessons Learned on Synthetic Data for Language Models
Wikipedia Bibliography: