The Scaling Hypothesis
https://www.reddit.com/r/mlscaling/comments/uznkhw/gpt3_2nd_anniversary/
https://www.reddit.com/r/mlscaling/comments/1fvsv8x/reviewing_the_2year_predictions_of_gpt3_2nd/
GPT-3: Language Models are Few-Shot Learners
Language Models are Unsupervised Multitask Learners
GPT-3: a Disappointing Paper
https://www.reddit.com/r/MachineLearning/comments/gteti8/d_gpt3_a_disappointing_paper/
OPT: Open Pre-trained Transformer Language Models
Recipes for building an open-domain chatbot
Galactica: A Large Language Model for Science
LLaMa-1: Open and Efficient Foundation Language Models
Gato: A Generalist Agent
Chinchilla: Training Compute-Optimal Large Language Models
Flamingo: a Visual Language Model for Few-Shot Learning
‘LaMDA’ directory
MUM: A New AI Milestone for Understanding Information
Scaling Language Models: Methods, Analysis & Insights from Training Gopher
PaLM: Scaling Language Modeling with Pathways
scaling-hypothesis#blessings-of-scale
[Transclude the forward-link's
context]
A Neural Algorithm of Artistic Style
GPT-3 Creative Fiction § Literary Parodies
[Transclude the forward-link's
context]
A Recipe For Arbitrary Text Style Transfer with Large Language Models
The Bitter Lesson
RNN Metadata for Mimicking Author Style
15.ai
Neonbjb/tortoise-Tts: A Multi-Voice TTS System Trained With an Emphasis on Quality
‘tech economics’ directory
‘Codex’ directory
Evaluating Large Language Models Trained on Code
https://github.com/features/copilot/
InstructGPT: Training language models to follow instructions with human feedback
Unsupervised Neural Machine Translation with Generative Language Models Only
index.md
https://openai.com/index/gpt-4-research/
‘video generation’ directory
Attention Is All You Need
https://suno.com/home
https://www.udio.com/
Smooth Adversarial Training
A Universal Law of Robustness via Isoperimetry
‘continual learning’ directory
‘tabular ML’ directory
scaling-hypothesis#why-does-pretraining-work
[Transclude the forward-link's
context]
Introducing Gemini Robotics and Gemini Robotics-ER, AI Models Designed for Robots to Understand, Act and React to the Physical World
Decision Transformer: Reinforcement Learning via Sequence Modeling
Perceiver: General Perception with Iterative Attention
Perceiver IO: A General Architecture for Structured Inputs & Outputs
Generating Diverse High-Fidelity Images with VQ-VAE-2
Timing Technology: Lessons From The Media Lab
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
PEER: Mixture of A Million Experts
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
‘MLP NN’ directory
MLP-Mixer: An all-MLP Architecture for Vision
WBE and DRL: a Middle Way of imitation learning
Accelerating progress in brain recording tech
CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers
PanGu-α: Large-scale Autoregressive Pretrained Chinese Language Models with Auto-parallel Computation
Chinese AI lab challenges Google, OpenAI with a model of 1.75 trillion parameters
China Has Already Reached Exascale—On Two Separate Systems
Google’s DeepMind-Brain merger: tech giant regroups for AI battle
https://openai.com/blog/chatgpt/
https://x.com/punk6529/status/1509832349986562048
Denoising Diffusion Probabilistic Models
https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2808734
‘experience curve’ directory
What Have The Romans…
https://x.com/bneyshabur/status/1529506103708602369
https://openai.com/index/dall-e-2/
https://openai.com/blog/reducing-bias-and-improving-safety-in-dall-e-2/
https://www.reddit.com/r/AnimeResearch/comments/txvu3a/anime_x_dalle_2_thread/
https://arxiv.org/pdf/2204.06125#page=16
GPT-3 Creative Fiction § BPEs
Critical Error in the Implementation of Compare_gan vs the Official BigGAN Model