Bibliography (26):

  1. A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play

  2. Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data

  3. Beyond Model Collapse: Scaling Up with Synthesized Data Requires Reinforcement

  4. How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse

  5. GPT-3: Language Models are Few-Shot Learners

  6. Deep reinforcement learning from human preferences

  7. Chinchilla: Training Compute-Optimal Large Language Models

  8. Grandmaster-Level Chess Without Search

  9. Towards a Human-like Open-Domain Chatbot

  10. Self-distillation: Born Again Neural Networks

  11. https://openai.com/o1/

  12. Grandmaster level in StarCraft II using multi-agent reinforcement learning

  13. ​ scaling-hypothesis#blessings-of-scale

    [Transclude the forward-link's context]

  14. ​ β€˜MARL’ directory

  15. Diversifying AI: Towards Creative Chess with AlphaZero (AZdb)

  16. Best Practices and Lessons Learned on Synthetic Data for Language Models

  17. LangChain Overview