Bibliography (105):

  1. The Scaling Hypothesis

  2. https://www.reddit.com/r/mlscaling/comments/uznkhw/gpt3_2nd_anniversary/

  3. https://www.reddit.com/r/mlscaling/comments/1fvsv8x/reviewing_the_2year_predictions_of_gpt3_2nd/

  4. GPT-3: Language Models are Few-Shot Learners

  5. Language Models are Unsupervised Multitask Learners

  6. GPT-3: a Disappointing Paper

  7. https://www.reddit.com/r/MachineLearning/comments/gteti8/d_gpt3_a_disappointing_paper/

  8. OPT: Open Pre-trained Transformer Language Models

  9. Recipes for building an open-domain chatbot

  10. Galactica: A Large Language Model for Science

  11. LLaMa-1: Open and Efficient Foundation Language Models

  12. Gato: A Generalist Agent

  13. Chinchilla: Training Compute-Optimal Large Language Models

  14. Flamingo: a Visual Language Model for Few-Shot Learning

  15. ‘LaMDA’ directory

  16. MUM: A New AI Milestone for Understanding Information

  17. Scaling Language Models: Methods, Analysis & Insights from Training Gopher

  18. PaLM: Scaling Language Modeling with Pathways

  19. scaling-hypothesis#blessings-of-scale

    [Transclude the forward-link's context]

  20. A Neural Algorithm of Artistic Style

  21. GPT-3 Creative Fiction § Literary Parodies

    [Transclude the forward-link's context]

  22. A Recipe For Arbitrary Text Style Transfer with Large Language Models

  23. The Bitter Lesson

  24. RNN Metadata for Mimicking Author Style

  25. 15.ai

  26. Neonbjb/tortoise-Tts: A Multi-Voice TTS System Trained With an Emphasis on Quality

  27. ‘tech economics’ directory

  28. ‘Codex’ directory

  29. Evaluating Large Language Models Trained on Code

  30. https://github.com/features/copilot/

  31. InstructGPT: Training language models to follow instructions with human feedback

  32. Unsupervised Neural Machine Translation with Generative Language Models Only

  33. index.md

  34. https://openai.com/index/gpt-4-research/

  35. ‘video generation’ directory

  36. Attention Is All You Need

  37. https://suno.com/home

  38. https://www.udio.com/

  39. Smooth Adversarial Training

  40. A Universal Law of Robustness via Isoperimetry

  41. ‘continual learning’ directory

  42. ‘tabular ML’ directory

  43. scaling-hypothesis#why-does-pretraining-work

    [Transclude the forward-link's context]

  44. Introducing Gemini Robotics and Gemini Robotics-ER, AI Models Designed for Robots to Understand, Act and React to the Physical World

  45. Decision Transformer: Reinforcement Learning via Sequence Modeling

  46. Perceiver: General Perception with Iterative Attention

  47. Perceiver IO: A General Architecture for Structured Inputs & Outputs

  48. Generating Diverse High-Fidelity Images with VQ-VAE-2

  49. Timing Technology: Lessons From The Media Lab

  50. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

  51. PEER: Mixture of A Million Experts

  52. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

  53. ‘MLP NN’ directory

  54. MLP-Mixer: An all-MLP Architecture for Vision

  55. WBE and DRL: a Middle Way of imitation learning

  56. Accelerating progress in brain recording tech

  57. CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

  58. PanGu-α: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation

  59. Chinese AI lab challenges Google, OpenAI with a model of 1.75 trillion parameters

  60. China Has Already Reached Exascale—On Two Separate Systems

  61. Google’s DeepMind-Brain merger: tech giant regroups for AI battle

  62. https://openai.com/blog/chatgpt/

  63. https://x.com/punk6529/status/1509832349986562048

  64. Denoising Diffusion Probabilistic Models

  65. https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2808734

  66. ‘experience curve’ directory

  67. What Have The Romans…

  68. https://x.com/bneyshabur/status/1529506103708602369

  69. https://openai.com/index/dall-e-2/

  70. https://openai.com/blog/reducing-bias-and-improving-safety-in-dall-e-2/

  71. https://www.reddit.com/r/AnimeResearch/comments/txvu3a/anime_x_dalle_2_thread/

  72. https://arxiv.org/pdf/2204.06125#page=16

  73. GPT-3 Creative Fiction § BPEs

  74. Critical Error in the Implementation of Compare_gan vs the Official BigGAN Model