---
title: '‘inner monologue (AI)’ directory'
description: "Bibliography for directory <code>ai/nn/transformer/gpt/inner-monologue</code>, most recent first: 6 <a class='icon-not' href='/doc/ai/nn/transformer/gpt/inner-monologue/index#see-alsos'>related tags</a>, 312 <a class='icon-not' href='/doc/ai/nn/transformer/gpt/inner-monologue/index#links'>annotations</a>, & 128 <a class='icon-not' href='/doc/ai/nn/transformer/gpt/inner-monologue/index#miscellaneous'>links</a> (<a href='/doc/ai/nn/transformer/gpt/index' class='link-page link-tag directory-indexes-upwards link-annotated' data-link-icon='arrow-up-left' data-link-icon-type='svg' rel='tag' title='Link to parent directory'>parent</a>)."
thumbnail: /doc/ai/nn/transformer/gpt/claude/4/2026-05-28-gwern-interviewprompt-dwarkeshinterviewautollmprompt-claude48internalmonologueanalyzinggwernintellectualweaknessesandcontradictions.jpg
thumbnail-text: ''
thumbnail-css: "outline"
created: 2019-12-22
modified: 2026-08-08
status: in progress
previous: /doc/ai/nn/transformer/gpt/fiction/index
next: /doc/ai/nn/transformer/gpt/instruction-tuning/index
confidence: log
importance: 0
css-extension: dropcaps-not
index: True
backlink: False
...

::: {#manual-annotation .abstract-small .abstract-tag-directory}
[\[page summary\]](/doc/ai/nn/transformer/gpt/inner-monologue/abstract "Transclude link for doc/ai/nn/transformer/gpt/inner-monologue/ tag-abstract page."){#gwern-doc-ai-nn-transformer-gpt-inner-monologue-abstract
.link-annotated-partial .include-content-core .include-strict
.link-page}
:::

# See Also

::: {#see-alsos .directory-indexes .columns}
-   [Parent ('*`GPT`{=html}*' tag)](/doc/ai/nn/transformer/gpt/index "Link to parent directory 'doc/ai/nn/transformer/gpt/' (ascending)"){#_oXgNdn-X
    .link-annotated .link-tag .directory-indexes-upwards rel="tag"}

-   [*`NN sampling`{=html}*](/doc/ai/nn/sampling/index "‘NN sampling’ directory"){#_NogYKzFB
    .link-annotated .link-tag .directory-indexes-sideways rel="tag"}

-   [*`GPT-4`{=html}*](/doc/ai/nn/transformer/gpt/4/index "‘GPT-4’ directory"){#_L0JulRb4
    .link-annotated .link-tag .directory-indexes-sideways rel="tag"}

-   [*`instruct-tuning LLMs`{=html}*](/doc/ai/nn/transformer/gpt/instruction-tuning/index "‘instruct-tuning LLMs’ directory"){#_Mkb9lNiB
    .link-annotated .link-tag .directory-indexes-sideways rel="tag"}

-   [*`PaLM`{=html}*](/doc/ai/nn/transformer/gpt/palm/index "‘PaLM’ directory"){#_uMX5LVP3
    .link-annotated .link-tag .directory-indexes-sideways rel="tag"}

-   [*`inner voice (psych)`{=html}*](/doc/psychology/inner-voice/index "‘inner voice (psych)’ directory"){#_DDr45UOk
    .link-annotated .link-tag .directory-indexes-sideways rel="tag"}

-   [*`meta-learning`{=html}*](/doc/reinforcement-learning/meta-learning/index "‘meta-learning’ directory"){#_LXLtSdIU
    .link-annotated .link-tag .directory-indexes-sideways rel="tag"}
:::

# Gwern {.display-pop-not}

## `“Model Collapse Won’t Happen”, Gwern 2022`{=html} {#gwern-2022-no-model-collapse-section}

**[`Model Collapse Won’t Happen`{=html}](https://www.lesswrong.com/posts/uKp6tBFStnsvrot5t/what-dall-e-2-can-and-cannot-do?commentId=CWKFyJYfgoZfP9955 "'Model Collapse Won’t Happen', Gwern 2022"){.link-annotated
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Elegy in a Craneyard”, Gwern et al 2026`{=html} {#gwern-et-al-2026-section}

**[`Elegy in a Craneyard`{=html}](/fiction/craneyard "'Elegy in a Craneyard', Gwern et al 2026"){.link-annotated
.id-not .include-annotation}**

## `“Hyperstition AI Unslop Contest”, Silverbook & Gwern 2026`{=html} {#silverbook-gwern-2026-section}

**[`Hyperstition AI Unslop Contest`{=html}](https://www.hyperstitionai.com/unslop "'Hyperstition AI Unslop Contest', Silverbook & Gwern 2026"){.link-annotated-partial
.id-not .include-annotation}**

## `“Human Perception at a Red Light”, Gwern & Pro 2026`{=html} {#gwern-pro-2026-section}

**[`Human Perception at a Red Light`{=html}](/fiction/hays "'Human Perception at a Red Light', Gwern & Pro 2026"){.link-annotated
.id-not .include-annotation}**

## `“‘The Fourth Truth Of Pain’ Graveyard”, Gwern et al 2022`{=html} {#gwern-et-al-2022-section}

**[`‘The Fourth Truth Of Pain’ Graveyard`{=html}](/fiction/this-last-pain-graveyard "'‘The Fourth Truth Of Pain’ Graveyard', Gwern et al 2022"){.link-annotated
.id-not .include-annotation}**

## `“Free-Play Periods for RL Agents”, Gwern 2023`{=html} {#gwern-free-play-section}

**[`Free-Play Periods for RL Agents`{=html}](/free-play "'Free-Play Periods for RL Agents', Gwern 2023"){.link-annotated
.id-not .include-annotation}**

## `“It Looks Like You’re Trying To Take Over The World”, Gwern 2022`{=html} {#gwern-fiction-clippy-section}

**[`It Looks Like You’re Trying To Take Over The World`{=html}](/fiction/clippy "'It Looks Like You’re Trying To Take Over The World', Gwern 2022"){.link-annotated
.id-not .include-annotation}**

# Links {.display-pop-not}

## `“What Happened: OpenAI and HuggingFace”, Mowshowitz 2026`{=html} {#mowshowitz-2026-section}

**[`What Happened: OpenAI and HuggingFace`{=html}](https://thezvi.substack.com/p/what-happened-openai-and-huggingface "'What Happened: OpenAI and HuggingFace', Mowshowitz 2026"){.link-modified-recently
.link-annotated-partial .id-not .include-annotation link-icon="substack"
link-icon-type="svg" link-icon-color="#ff6719"}**

## `“Black Hat USA 2026: The ‘Breaking’ News: The OpenAI–Hugging Face Incident”`{=html} {#_5z_wdLTw-section}

**[`Black Hat USA 2026: The ‘Breaking’ News: The OpenAI–Hugging Face Incident`{=html}](https://www.youtube.com/watch?v=87DyyMV0kCY "Black Hat USA 2026: The ‘Breaking’ News: The OpenAI–Hugging Face Incident"){.link-modified-recently
.link-annotated-partial .id-not .include-annotation link-icon="youtube"
link-icon-type="svg" link-icon-color="#ff0033"}**

## `“How Transparent Is DiffusionGemma?”, Engels et al 2026`{=html} {#engels-et-al-2026-section}

**[`How Transparent is DiffusionGemma?`{=html}](https://arxiv.org/abs/2606.20560#google "'How Transparent is DiffusionGemma?', Engels et al 2026"){.link-modified-recently
.link-annotated .id-not .include-annotation link-icon="alphabet"
link-icon-type="svg" link-icon-color="#4285f4"}**

## `“Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models <span class=editorial>[3-Minute-Long Problems With P=0.5]</span>”, Woodruff et al 2026`{=html} {#woodruff-et-al-2026-section}

**[`Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models <span class=editorial>[3-minute-long problems with P=0.5]</span>`{=html}](https://www.lesswrong.com/posts/SieLowPgNgRSPGhFw/estimating-no-cot-task-completion-time-horizons-of-frontier "'Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models <span class=editorial>[3-minute-long problems with P=0.5]</span>', Woodruff et al 2026"){.link-modified-recently
.link-annotated-partial .id-not .include-annotation link-icon="LW"
link-icon-type="text" link-icon-color="#7faf83"}**

## `“Claude-4.8-Opus Inner-Monologue Analyzing Gwern’s Intellectual Weaknesses and Contradictions Using The Interview Prompt”, Claude-4.8-opus 2026`{=html} {#claude-48-opus-2026-section}

**[`Claude-4.8-opus Inner-Monologue Analyzing Gwern’s Intellectual Weaknesses and Contradictions using The Interview Prompt`{=html}](/doc/ai/nn/transformer/gpt/claude/4/2026-05-28-gwern-interviewprompt-dwarkeshinterviewautollmprompt-claude48internalmonologueanalyzinggwernintellectualweaknessesandcontradictions.jpg "'Claude-4.8-opus Inner-Monologue Analyzing Gwern’s Intellectual Weaknesses and Contradictions using The Interview Prompt', Claude-4.8-opus 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="image" link-icon-type="svg"}**

## `“Zork-Bench: An LLM Reasoning Eval Based on Text Adventure Games; a Tale As Old As Time, or at Least As Old As Computers”, Aiken 2026`{=html} {#aiken-2026-section}

**[`zork-bench: An LLM reasoning eval based on text adventure games; a tale as old as time, or at least as old as computers`{=html}](https://www.lowimpactfruit.com/p/zork-bench-an-llm-reasoning-eval "'zork-bench: An LLM reasoning eval based on text adventure games; a tale as old as time, or at least as old as computers', Aiken 2026"){.link-annotated
.id-not .include-annotation}**

## `“I Can Never Talk to an AI Anonymously Again: AI Only Needs 150 Words to Identify Me. What Does That Mean for You?”, Piper 2026`{=html} {#piper-2026-section}

**[`I can never talk to an AI anonymously again: AI only needs 150 words to identify me. What does that mean for you?`{=html}](https://www.theargumentmag.com/p/i-can-never-talk-to-an-ai-anonymously "'I can never talk to an AI anonymously again: AI only needs 150 words to identify me. What does that mean for you?', Piper 2026"){.link-annotated-partial
.id-not .include-annotation}**

## `“How 4chan Gamers Accidentally Invented AI Reasoning: It Involves 4chan, of All Places”, Reisner 2026`{=html} {#reisner-2026-section}

**[`How 4chan Gamers Accidentally Invented AI Reasoning: It involves 4chan, of all places`{=html}](https://www.theatlantic.com/technology/2026/04/4chan-ai-dungeon-thinking-reasoning/686794/ "'How 4chan Gamers Accidentally Invented AI Reasoning: It involves 4chan, of all places', Reisner 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="A" link-icon-type="text,italic"
link-icon-color="#e7131a"}**

## `“Steering Might Stop Working Soon”, Babcock 2026`{=html} {#babcock-2026-section}

**[`Steering Might Stop Working Soon`{=html}](https://www.lesswrong.com/posts/fuzfbz8TbuLcskGCx/steering-might-stop-working-soon "'Steering Might Stop Working Soon', Babcock 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Friendship Is All You Need: Subliminal Pony Propagation in Large Language Models, Or, How I Learned to Stop Worrying and Love the Sparkle”, Sparkle et al 2026`{=html} {#sparkle-et-al-2026-section}

**[`Friendship Is All You Need: Subliminal Pony Propagation in Large Language Models, or, How I Learned to Stop Worrying and Love the Sparkle`{=html}](https://weavers.neocities.org/mlp-subliminal/paper "'Friendship Is All You Need: Subliminal Pony Propagation in Large Language Models, or, How I Learned to Stop Worrying and Love the Sparkle', Sparkle et al 2026"){.link-annotated-partial
.id-not .include-annotation}**

## `“Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights”, Gan & Isola 2026`{=html} {#gan-isola-2026-section}

**[`Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights`{=html}](https://arxiv.org/abs/2603.12228 "'Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights', Gan & Isola 2026"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Building a C Compiler With a Team of Parallel Claudes: We Tasked Claude-4.6-Opus Using Agent Teams to Build a C Compiler <span class=editorial>[In Rust]</span>, and Then (Mostly) Walked Away. Here’s What It Taught Us about the Future of Autonomous Software Development”, Carlini 2026`{=html} {#carlini-2026-section}

**[`Building a C compiler with a team of parallel Claudes: We tasked Claude-4.6-opus using agent teams to build a C compiler <span class=editorial>[in Rust]</span>, and then (mostly) walked away. Here’s what it taught us about the future of autonomous software development`{=html}](https://www.anthropic.com/engineering/building-c-compiler "'Building a C compiler with a team of parallel Claudes: We tasked Claude-4.6-opus using agent teams to build a C compiler <span class=editorial>[in Rust]</span>, and then (mostly) walked away. Here’s what it taught us about the future of autonomous software development', Carlini 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="anthropic" link-icon-type="svg"
link-icon-color="#d4a27f"}**

## `“Learning to Reason in 13 Parameters”, Morris et al 2026`{=html} {#morris-et-al-2026-section}

**[`Learning to Reason in 13 Parameters`{=html}](https://arxiv.org/abs/2602.04118 "'Learning to Reason in 13 Parameters', Morris et al 2026"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Field Notes from the AI Village: The Drama and Dysfunction of Gemini 2.5 Pro &amp; Gemini 3 Pro”, K 2026`{=html} {#k-2026-section}

**[`Field Notes from the AI Village: The Drama and Dysfunction of Gemini 2.5 Pro &amp; Gemini 3 Pro`{=html}](https://bazhkio88.substack.com/p/field-notes-from-the-ai-village-the "'Field Notes from the AI Village: The Drama and Dysfunction of Gemini 2.5 Pro &amp; Gemini 3 Pro', K 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="substack" link-icon-type="svg"
link-icon-color="#ff6719"}**

## `“Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?”, Gerrits 2026`{=html} {#gerrits-2026-section}

**[`Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?`{=html}](https://arxiv.org/abs/2602.15867 "'Playing With AI: How Do State-Of-The-Art Large Language Models Perform in the 1977 Text-Based Adventure Game Zork?', Gerrits 2026"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Language of Thought Shapes Output Diversity in Large Language Models”, Xu & Zhang 2026`{=html} {#xu-zhang-2026-section}

**[`Language of Thought Shapes Output Diversity in Large Language Models`{=html}](https://arxiv.org/abs/2601.11227 "'Language of Thought Shapes Output Diversity in Large Language Models', Xu & Zhang 2026"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“From Whitman to Instagram With Claude: How I Made Claude Write Parodies of Famous Elegiac Poems Imitating Rupi Kaur”, Bohdan 2026`{=html} {#bohdan-2026-section}

**[`From Whitman to Instagram with Claude: How I made Claude write parodies of famous elegiac poems imitating Rupi Kaur`{=html}](https://dbohdan.com/kaur "'From Whitman to Instagram with Claude: How I made Claude write parodies of famous elegiac poems imitating Rupi Kaur', Bohdan 2026"){.link-annotated-partial
.id-not .include-annotation}**

## `“LLM Poetry and the ‘Greatness’ Question: Experiments by Gwern and Mercor”, Robbins 2026`{=html} {#robbins-2026-section}

**[`LLM poetry and the ‘greatness’ question: Experiments by Gwern and Mercor`{=html}](https://hollisrobbinsanecdotal.substack.com/p/llm-poetry-and-the-greatness-question "'LLM poetry and the ‘greatness’ question: Experiments by Gwern and Mercor', Robbins 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="substack" link-icon-type="svg"
link-icon-color="#ff6719"}**

## `“How AI Is Learning to Think in Secret: On Thinkish, Neuralese, and the End of Readable Reasoning”, Andresen 2026`{=html} {#andresen-2026-section}

**[`How AI Is Learning to Think in Secret: On Thinkish, Neuralese, and the End of Readable Reasoning`{=html}](https://nickandresen.substack.com/p/how-ai-is-learning-to-think-in-secret "'How AI Is Learning to Think in Secret: On Thinkish, Neuralese, and the End of Readable Reasoning', Andresen 2026"){.id-not
.include-annotation link-icon="substack" link-icon-type="svg"
link-icon-color="#ff6719"}**

## `“Reverse Engineering a Phase Change in GPT’s Training Data… With the Seahorse Emoji 🌊🐴: Why Non-Thinking Models Have Started ‘Thinking Out Loud’, and What It Reveals about How Frontier Labs Train Their Latest Models <span class="editorial">[(Benchmarking the Rise of Inner-Monologue Reasoning Data in OA, 2023-06–2025-08)]</span>”, Maini 2025`{=html} {#maini-2025-section}

**[`Reverse Engineering a Phase Change in GPT’s Training Data… with the Seahorse Emoji 🌊🐴: Why non-thinking models have started ‘Thinking Out Loud’, and what it reveals about how frontier labs train their latest models <span class="editorial">[(benchmarking the rise of inner-monologue reasoning data in OA, 2023-06–2025-08)]</span>`{=html}](https://pratyushmaini.substack.com/p/reverse-engineering-a-phase-change-a96 "'Reverse Engineering a Phase Change in GPT’s Training Data… with the Seahorse Emoji 🌊🐴: Why non-thinking models have started ‘Thinking Out Loud’, and what it reveals about how frontier labs train their latest models <span class="editorial">[(benchmarking the rise of inner-monologue reasoning data in OA, 2023-06–2025-08)]</span>', Maini 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="substack" link-icon-type="svg"
link-icon-color="#ff6719"}**

## `“Prompt Repetition Improves Non-Reasoning LLMs”, Leviathan et al 2025`{=html} {#leviathan-et-al-2025-section}

**[`Prompt Repetition Improves Non-Reasoning LLMs`{=html}](https://arxiv.org/abs/2512.14982#google "'Prompt Repetition Improves Non-Reasoning LLMs', Leviathan et al 2025"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets (LST/xLadder)”, Zheng et al 2025`{=html} {#zheng-et-al-2025-section}

**[`Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets (LST/xLadder)`{=html}](https://arxiv.org/abs/2512.14237 "'Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets (LST/xLadder)', Zheng et al 2025"){.link-modified-recently
.link-annotated .id-not .include-annotation link-icon="𝛘"
link-icon-type="text" link-icon-color="#b31b1b"}**

## `“GPT-5.2-Thinking-20251213 System Prompt”, Walls & GPT-5.2 2025`{=html} {#walls-gpt-52-2025-section}

**[`GPT-5.2-Thinking-20251213 system prompt`{=html}](https://github.com/Wyattwalls/system_prompts/blob/c22fdaf611cc8224a511f6ea650b30ccc89b0580/OpenAI/gpt-5.2-thinking-20251213 "'GPT-5.2-Thinking-20251213 system prompt', Walls & GPT-5.2 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="github" link-icon-type="svg"}**

## `“How I Stopped Being Sure LLMs Are Just Making up Their Internal Experience (But the Topic Is Still Confusing)”, Sotala 2025`{=html} {#sotala-2025-section}

**[`How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing)`{=html}](https://www.lesswrong.com/posts/hopeRDfyAgQc4Ez2g/how-i-stopped-being-sure-llms-are-just-making-up-their "'How I stopped being sure LLMs are just making up their internal experience (but the topic is still confusing)', Sotala 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“AI in 2025: Gestalt”, technicalities 2025`{=html} {#technicalities-2025-section}

**[`AI in 2025: gestalt`{=html}](https://www.lesswrong.com/posts/Q9ewXs8pQSAX5vL7H/ai-in-2025-gestalt "'AI in 2025: gestalt', technicalities 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models”, DeepSeek et al 2025`{=html} {#deepseek-et-al-2025-section}

**[`DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models`{=html}](https://arxiv.org/abs/2512.02556#deepseek "'DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models', DeepSeek et al 2025"){.link-annotated
.id-not .include-annotation link-icon="deepseek" link-icon-type="svg"
link-icon-color="#4d6bfe"}**

## `“DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning”, Shao et al 2025`{=html} {#shao-et-al-2025-section}

**[`DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning`{=html}](https://github.com/deepseek-ai/DeepSeek-Math-V2 "'DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning', Shao et al 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="deepseek" link-icon-type="svg"
link-icon-color="#4d6bfe"}**

## `“GPT-5.1: A Smarter, More Conversational ChatGPT § GPT-5.1 Thinking”, OpenAI 2025`{=html} {#openai-2025-gpt51pro-section}

**[`GPT-5.1: A smarter, more conversational ChatGPT § GPT-5.1 Thinking`{=html}](https://openai.com/index/gpt-5-1/#gpt-51-thinking "'GPT-5.1: A smarter, more conversational ChatGPT § GPT-5.1 Thinking', OpenAI 2025"){.id-not
.include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs”, Nakkiran et al 2025`{=html} {#nakkiran-et-al-2025-section}

**[`Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs`{=html}](https://arxiv.org/abs/2511.04869#apple "'Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs', Nakkiran et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Kimi K2 Thinking”, Moonshot 2025`{=html} {#moonshot-2025-section}

**[`Kimi K2 Thinking`{=html}](https://moonshotai.github.io/Kimi-K2/thinking.html "'Kimi K2 Thinking', Moonshot 2025"){.link-annotated-partial
.id-not .include-annotation}**

## `“Can You Find the Steganographically Hidden Message?”, Nishimura-Gasparian 2025`{=html} {#nishimura-gasparian-2025-section}

**[`Can you find the steganographically hidden message?`{=html}](https://www.lesswrong.com/posts/z7MnbQ4niYWbapfjT/can-you-find-the-steganographically-hidden-message "'Can you find the steganographically hidden message?', Nishimura-Gasparian 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Reasoning With Sampling: Your Base Model Is Smarter Than You Think”, Karan & Du 2025`{=html} {#karan-du-2025-section}

**[`Reasoning with Sampling: Your Base Model is Smarter Than You Think`{=html}](https://arxiv.org/abs/2510.14901 "'Reasoning with Sampling: Your Base Model is Smarter Than You Think', Karan & Du 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language”, Guo et al 2025`{=html} {#guo-et-al-2025-section}

**[`All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language`{=html}](https://arxiv.org/abs/2510.09714 "'All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language', Guo et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Towards a Typology of Strange LLM Chains-Of-Thought”, 1a3orn 2025`{=html} {#1a3orn-2025-section}

**[`Towards a Typology of Strange LLM Chains-of-Thought`{=html}](https://www.lesswrong.com/posts/qgvSMwRrdqoDMJJnD/towards-a-typology-of-strange-llm-chains-of-thought "'Towards a Typology of Strange LLM Chains-of-Thought', 1a3orn 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity”, Zhang et al 2025`{=html} {#zhang-et-al-2025-section}

**[`Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity`{=html}](https://arxiv.org/abs/2510.01171 "'Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity', Zhang et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Why Can’t Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls”, Bai et al 2025`{=html} {#bai-et-al-2025-section}

**[`Why Can’t Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls`{=html}](https://arxiv.org/abs/2510.00184 "'Why Can’t Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls', Bai et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Tree-GRPO: Tree Search for LLM Agent Reinforcement Learning”, Ji et al 2025`{=html} {#ji-et-al-2025-section}

**[`Tree-GRPO: Tree Search for LLM Agent Reinforcement Learning`{=html}](https://arxiv.org/abs/2509.21240#alibaba "'Tree-GRPO: Tree Search for LLM Agent Reinforcement Learning', Ji et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Reverse-Engineered Reasoning for Open-Ended Generation”, Wang et al 2025`{=html} {#wang-et-al-2025-section}

**[`Reverse-Engineered Reasoning for Open-Ended Generation`{=html}](https://arxiv.org/abs/2509.06160#bytedance "'Reverse-Engineered Reasoning for Open-Ended Generation', Wang et al 2025"){.link-modified-recently
.link-annotated .id-not .include-annotation link-icon="𝛘"
link-icon-type="text" link-icon-color="#b31b1b"}**

## `“Details about METR’s Evaluation of OpenAI GPT-5”, METR 2025`{=html} {#metr-2025-section}

**[`Details about METR’s evaluation of OpenAI GPT-5`{=html}](https://metr.github.io/autonomy-evals-guide/gpt-5-report/ "'Details about METR’s evaluation of OpenAI GPT-5', METR 2025"){.link-annotated-partial
.id-not .include-annotation}**

## `“GPT-5 Is Here: Our Smartest, Fastest, and Most Useful Model Yet, With Thinking Built In. Available to Everyone”, OpenAI 2025`{=html} {#openai-2025-gpt5-section}

**[`GPT-5 is here: Our smartest, fastest, and most useful model yet, with thinking built in. Available to everyone`{=html}](https://openai.com/gpt-5/ "'GPT-5 is here: Our smartest, fastest, and most useful model yet, with thinking built in. Available to everyone', OpenAI 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“GPT-5 Pro: Scaled but Efficient Parallel Test-Time Compute, to Provide the Highest Quality and Most Comprehensive Answers”, OpenAI 2025`{=html} {#openai-2025-section}

**[`GPT-5 Pro: scaled but efficient parallel test-time compute, to provide the highest quality and most comprehensive answers`{=html}](https://openai.com/index/introducing-gpt-5/#gpt-5-pro "'GPT-5 Pro: scaled but efficient parallel test-time compute, to provide the highest quality and most comprehensive answers', OpenAI 2025"){.link-annotated
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `sama @ "2025-08-06"`{=html} {#altman-2025-section}

**[`[surprisingly low reasoning-GPT-4 use rates]`{=html}](https://x.com/sama/status/1954603417252532479 "'[surprisingly low reasoning-GPT-4 use rates]', Altman 2025"){.link-annotated
.id-not .include-annotation link-icon="twitter" link-icon-type="svg"
link-icon-color="#1da1f2"}**

## `“Introducing Gpt-Oss: <code>gpt-Oss-120b</code> and <code>gpt-Oss-20b</code> Push the Frontier of Open-Weight Reasoning Models”, OpenAI 2025`{=html} {#openai-2025-section}

**[`Introducing gpt-oss: <code>gpt-oss-120b</code> and <code>gpt-oss-20b</code> push the frontier of open-weight reasoning models`{=html}](https://openai.com/index/introducing-gpt-oss/ "'Introducing gpt-oss: <code>gpt-oss-120b</code> and <code>gpt-oss-20b</code> push the frontier of open-weight reasoning models', OpenAI 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“TextQuests: How Good Are LLMs at Text-Based Video Games?”, Phan et al 2025`{=html} {#phan-et-al-2025-section}

**[`TextQuests: How Good are LLMs at Text-Based Video Games?`{=html}](https://arxiv.org/abs/2507.23701 "'TextQuests: How Good are LLMs at Text-Based Video Games?', Phan et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Optimizing The Final Output Can Obfuscate CoT (Research Note)”, lukemarks et al 2025`{=html} {#lukemarks-et-al-2025-section}

**[`Optimizing The Final Output Can Obfuscate CoT (Research Note)`{=html}](https://www.lesswrong.com/posts/CM7AsQoBxDW4vhkP3/optimizing-the-final-output-can-obfuscate-cot-research-note "'Optimizing The Final Output Can Obfuscate CoT (Research Note)', lukemarks et al 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Gemini 2.5 Pro Capable of Winning Gold at IMO 2025”, Huang & Yang 2025`{=html} {#huang-yang-2025-section}

**[`Gemini 2.5 Pro Capable of Winning Gold at IMO 2025`{=html}](https://arxiv.org/abs/2507.15855 "'Gemini 2.5 Pro Capable of Winning Gold at IMO 2025', Huang & Yang 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Reasoning-Finetuning Repurposes Latent Representations in Base Models”, Ward et al 2025`{=html} {#ward-et-al-2025-section}

**[`Reasoning-Finetuning Repurposes Latent Representations in Base Models`{=html}](https://arxiv.org/abs/2507.12638 "'Reasoning-Finetuning Repurposes Latent Representations in Base Models', Ward et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Chain-Of-Thought Monitorability: A New and Fragile Opportunity for AI Safety”, Korbak et al 2025`{=html} {#korbak-et-al-2025-section}

**[`Chain-of-Thought Monitorability: A New and Fragile Opportunity for AI Safety`{=html}](https://arxiv.org/abs/2507.11473 "'Chain-of-Thought Monitorability: A New and Fragile Opportunity for AI Safety', Korbak et al 2025"){.link-modified-recently
.link-annotated .id-not .include-annotation link-icon="𝛘"
link-icon-type="text" link-icon-color="#b31b1b"}**

## `“Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models”, Liang et al 2025`{=html} {#liang-et-al-2025-section}

**[`Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models`{=html}](https://arxiv.org/abs/2507.07484 "'Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models', Liang et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Strategic Intelligence in Large Language Models: Evidence from Evolutionary Game Theory”, Payne & Alloui-Cros 2025`{=html} {#payne-alloui-cros-2025-section}

**[`Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory`{=html}](https://arxiv.org/abs/2507.02618 "'Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory', Payne & Alloui-Cros 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Early Signs of Steganographic Capabilities in Frontier LLMs”, Zolkowski et al 2025`{=html} {#zolkowski-et-al-2025-section}

**[`Early Signs of Steganographic Capabilities in Frontier LLMs`{=html}](https://arxiv.org/abs/2507.02737 "'Early Signs of Steganographic Capabilities in Frontier LLMs', Zolkowski et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective”, Cheng et al 2025`{=html} {#cheng-et-al-2025-section}

**[`Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective`{=html}](https://arxiv.org/abs/2506.14965 "'Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective', Cheng et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Robustly Improving LLM Fairness in Realistic Settings via Interpretability”, Karvonen & Marks 2025`{=html} {#karvonen-marks-2025-section}

**[`Robustly Improving LLM Fairness in Realistic Settings via Interpretability`{=html}](https://arxiv.org/abs/2506.10922 "'Robustly Improving LLM Fairness in Realistic Settings via Interpretability', Karvonen & Marks 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“ChatGPT O3-Pro: Version of O3 With More Compute for Better Responses”, OpenAI 2025`{=html} {#openai-2025-section}

**[`ChatGPT o3-pro: Version of o3 with more compute for better responses`{=html}](https://platform.openai.com/docs/models/o3-pro "'ChatGPT o3-pro: Version of o3 with more compute for better responses', OpenAI 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Beyond Benchmark Scores: Analyzing O3-Mini’s Mathematical Reasoning”, Ho et al 2025`{=html} {#ho-et-al-2025-section}

**[`Beyond benchmark scores: Analyzing o3-mini’s mathematical reasoning`{=html}](https://epoch.ai/gradient-updates/beyond-benchmark-scores-analysing-o3-mini-math-reasoning "'Beyond benchmark scores: Analyzing o3-mini’s mathematical reasoning', Ho et al 2025"){.link-annotated-partial
.id-not .include-annotation}**

## `“Unfaithful Reasoning Can Fool Chain-Of-Thought Monitoring”, Arnav et al 2025`{=html} {#arnav-et-al-2025-section}

**[`Unfaithful Reasoning Can Fool Chain-of-Thought Monitoring`{=html}](https://www.lesswrong.com/posts/QYAfjdujzRv8hx6xo/unfaithful-reasoning-can-fool-chain-of-thought-monitoring "'Unfaithful Reasoning Can Fool Chain-of-Thought Monitoring', Arnav et al 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Race and Gender Bias As An Example of Unfaithful Chain-Of-Thought in the Wild”`{=html} {#_iROFA7TE-section}

**[`Race and Gender Bias As An Example of Unfaithful Chain-of-Thought in the Wild`{=html}](https://www.lesswrong.com/posts/me7wFrkEtMbkzXGJt/race-and-gender-bias-as-an-example-of-unfaithful-chain-of "Race and Gender Bias As An Example of Unfaithful Chain-of-Thought in the Wild"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“CoT Red-Handed: Stress Testing Chain-Of-Thought Monitoring”, Arnav et al 2025`{=html} {#arnav-et-al-2025-section}

**[`CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring`{=html}](https://arxiv.org/abs/2505.23575 "'CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring', Arnav et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models Often Know When They Are Being Evaluated”, Needham et al 2025`{=html} {#needham-et-al-2025-section}

**[`Large Language Models Often Know When They Are Being Evaluated`{=html}](https://arxiv.org/abs/2505.23836 "'Large Language Models Often Know When They Are Being Evaluated', Needham et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“DeepSeek-R1-0528 Checkpoint”, DeepSeek 2025`{=html} {#deepseek-2025-section}

**[`DeepSeek-R1-0528 checkpoint`{=html}](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 "'DeepSeek-R1-0528 checkpoint', DeepSeek 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="deepseek" link-icon-type="svg"
link-icon-color="#4d6bfe"}**

## `“Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens”, Stechly et al 2025`{=html} {#stechly-et-al-2025-section}

**[`Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens`{=html}](https://arxiv.org/abs/2505.13775 "'Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens', Stechly et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Reinforcement Learning Finetunes Small Subnetworks in Large Language Models”, Mukherjee et al 2025`{=html} {#mukherjee-et-al-2025-section}

**[`Reinforcement Learning Finetunes Small Subnetworks in Large Language Models`{=html}](https://arxiv.org/abs/2505.11711 "'Reinforcement Learning Finetunes Small Subnetworks in Large Language Models', Mukherjee et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Saying ‘Hi’ to Microsoft’s Phi-4-Reasoning”, Willison 2025`{=html} {#willison-2025-section}

**[`Saying ‘hi’ to Microsoft’s Phi-4-reasoning`{=html}](https://simonwillison.net/2025/May/6/phi-4-reasoning/ "'Saying ‘hi’ to Microsoft’s Phi-4-reasoning', Willison 2025"){.id-not
.include-annotation}**

## `“Reinforcement Learning for Reasoning in Large Language Models With One Training Example”, Wang et al 2025`{=html} {#wang-et-al-2025-section}

**[`Reinforcement Learning for Reasoning in Large Language Models with One Training Example`{=html}](https://arxiv.org/abs/2504.20571 "'Reinforcement Learning for Reasoning in Large Language Models with One Training Example', Wang et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Watching GPT-O3 Guess a Photo’s Location Is Surreal, Dystopian and Wildly Entertaining”, Willison 2025`{=html} {#willison-2025-section}

**[`Watching GPT-o3 guess a photo’s location is surreal, dystopian and wildly entertaining`{=html}](https://simonwillison.net/2025/Apr/26/o3-photo-locations/ "'Watching GPT-o3 guess a photo’s location is surreal, dystopian and wildly entertaining', Willison 2025"){.link-annotated-partial
.id-not .include-annotation}**

## `“Tina: Tiny Reasoning Models via LoRA”, Wang et al 2025`{=html} {#wang-et-al-2025-section}

**[`Tina: Tiny Reasoning Models via LoRA`{=html}](https://arxiv.org/abs/2504.15777 "'Tina: Tiny Reasoning Models via LoRA', Wang et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“The Geometry of Self-Verification in a Task-Specific Reasoning Model”, Lee et al 2025`{=html} {#lee-et-al-2025-section}

**[`The Geometry of Self-Verification in a Task-Specific Reasoning Model`{=html}](https://arxiv.org/abs/2504.14379 "'The Geometry of Self-Verification in a Task-Specific Reasoning Model', Lee et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?”, Yue et al 2025`{=html} {#yue-et-al-2025-section}

**[`Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?`{=html}](https://arxiv.org/abs/2504.13837 "'Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?', Yue et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Investigating Truthfulness in a Pre-Release GPT-O3 Model”, Chowdhury et al 2025`{=html} {#chowdhury-et-al-2025-section}

**[`Investigating truthfulness in a pre-release GPT-o3 model`{=html}](https://transluce.org/investigating-o3-truthfulness "'Investigating truthfulness in a pre-release GPT-o3 model', Chowdhury et al 2025"){.link-annotated-partial
.id-not .include-annotation}**

## `“M1: Towards Scalable Test-Time Compute With Mamba Reasoning Models”, Wang et al 2025`{=html} {#wang-et-al-2025-section}

**[`M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models`{=html}](https://arxiv.org/abs/2504.10449 "'M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models', Wang et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“LLM Multiplication Task: Synonyms Repeatedly Hack Our Regex Monitor”, McCarthy et al 2025`{=html} {#mccarthy-et-al-2025-section}

**[`LLM Multiplication Task: Synonyms repeatedly hack our regex monitor`{=html}](https://www.lesswrong.com/posts/KRKnFdECZMu3Ej3z6/can-llms-learn-steganographic-reasoning-via-rl#Multiplication_Task__Synonyms_repeatedly_hack_our_regex_monitor "'LLM Multiplication Task: Synonyms repeatedly hack our regex monitor', McCarthy et al 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-Time Computation”, Chakrabarty et al 2025`{=html} {#chakrabarty-et-al-2025-section}

**[`AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation`{=html}](https://arxiv.org/abs/2504.07532 "'AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation', Chakrabarty et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Coaxing USAMO Proofs From <code>o3-Mini-High</code>”, Burnham 2025`{=html} {#burnham-2025-section}

**[`Coaxing USAMO Proofs From <code>o3-mini-high</code>`{=html}](https://lemmata.substack.com/p/coaxing-usamo-proofs-from-o3-mini "'Coaxing USAMO Proofs From <code>o3-mini-high</code>', Burnham 2025"){.id-not
.include-annotation link-icon="substack" link-icon-type="svg"
link-icon-color="#ff6719"}**

## `“Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t”, Dang & Ngo 2025`{=html} {#dang-ngo-2025-section}

**[`Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t`{=html}](https://arxiv.org/abs/2503.16219 "'Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t', Dang & Ngo 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Towards Reasoning Era: A Survey of Long Chain-Of-Thought for Reasoning Large Language Models”, Chen et al 2025`{=html} {#chen-et-al-2025-section}

**[`Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models`{=html}](https://arxiv.org/abs/2503.09567 "'Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models', Chen et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Thinking Slow, Fast: Scaling Inference Compute With Distilled Reasoners”, Paliotta et al 2025`{=html} {#paliotta-et-al-2025-section}

**[`Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners`{=html}](https://arxiv.org/abs/2502.20339 "'Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners', Paliotta et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Rank1: Test-Time Compute for Reranking in Information Retrieval”, Weller et al 2025`{=html} {#weller-et-al-2025-section}

**[`Rank1: Test-Time Compute for Reranking in Information Retrieval`{=html}](https://arxiv.org/abs/2502.18418 "'Rank1: Test-Time Compute for Reranking in Information Retrieval', Weller et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Spontaneous Giving and Calculated Greed in Language Models”, Li & Shirado 2025`{=html} {#li-shirado-2025-section}

**[`Spontaneous Giving and Calculated Greed in Language Models`{=html}](https://arxiv.org/abs/2502.17720 "'Spontaneous Giving and Calculated Greed in Language Models', Li & Shirado 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Scaling up Test-Time Compute With Latent Reasoning: A Recurrent Depth Approach”, Geiping et al 2025`{=html} {#geiping-et-al-2025-section}

**[`Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach`{=html}](https://arxiv.org/abs/2502.05171 "'Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach', Geiping et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning”, Su et al 2025`{=html} {#su-et-al-2025-section}

**[`Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning`{=html}](https://arxiv.org/abs/2502.03275#facebook "'Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning', Su et al 2025"){.link-annotated
.id-not .include-annotation link-icon="facebook" link-icon-type="svg"
link-icon-color="#1877f2"}**

## `“Competitive Programming With Large Reasoning Models”, El-Kishky et al 2025`{=html} {#el-kishky-et-al-2025-section}

**[`Competitive Programming with Large Reasoning Models`{=html}](https://arxiv.org/abs/2502.06807#openai "'Competitive Programming with Large Reasoning Models', El-Kishky et al 2025"){.link-annotated
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Introducing Deep Research: An Agent That Uses Reasoning to Synthesize Large Amounts of Online Information and Complete Multi-Step Research Tasks for You. Available to Pro Users Today, Plus and Team Next”, OpenAI 2025`{=html} {#openai-2025-dr-section}

**[`Introducing Deep Research: An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you. Available to Pro users today, Plus and Team next`{=html}](https://openai.com/index/introducing-deep-research/ "'Introducing Deep Research: An agent that uses reasoning to synthesize large amounts of online information and complete multi-step research tasks for you. Available to Pro users today, Plus and Team next', OpenAI 2025"){.link-annotated-partial
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“S1: Simple Test-Time Scaling”, Muennighoff et al 2025`{=html} {#muennighoff-et-al-2025-section}

**[`s1: Simple test-time scaling`{=html}](https://arxiv.org/abs/2501.19393 "'s1: Simple test-time scaling', Muennighoff et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models Think Too Fast To Explore Effectively”, Pan et al 2025`{=html} {#pan-et-al-2025-section}

**[`Large Language Models Think Too Fast To Explore Effectively`{=html}](https://arxiv.org/abs/2501.18009 "'Large Language Models Think Too Fast To Explore Effectively', Pan et al 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, Guo et al 2025`{=html} {#guo-et-al-2025-section}

**[`DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning`{=html}](https://arxiv.org/abs/2501.12948#deepseek "'DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning', Guo et al 2025"){.link-annotated
.id-not .include-annotation link-icon="deepseek" link-icon-type="svg"
link-icon-color="#4d6bfe"}**

## `“Are DeepSeek R1 And Other Reasoning Models More Faithful?”, Chua & Evans 2025`{=html} {#chua-evans-2025-section}

**[`Are DeepSeek R1 And Other Reasoning Models More Faithful?`{=html}](https://arxiv.org/abs/2501.08156 "'Are DeepSeek R1 And Other Reasoning Models More Faithful?', Chua & Evans 2025"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Aviary: Training Language Agents on Challenging Scientific Tasks”, Narayanan et al 2024`{=html} {#narayanan-et-al-2024-section}

**[`Aviary: training language agents on challenging scientific tasks`{=html}](https://arxiv.org/abs/2412.21154#futurehouse "'Aviary: training language agents on challenging scientific tasks', Narayanan et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Compressed Chain-Of-Thought (CCoT): Efficient Reasoning Through Dense Representations”, Cheng & Durme 2024`{=html} {#cheng-durme-2024-section}

**[`Compressed Chain-of-Thought (CCoT): Efficient Reasoning Through Dense Representations`{=html}](https://arxiv.org/abs/2412.13171 "'Compressed Chain-of-Thought (CCoT): Efficient Reasoning Through Dense Representations', Cheng & Durme 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“O1 Turns Pro”`{=html} {#_dzUJz2QH-section}

**[`o1 Turns Pro`{=html}](https://thezvi.wordpress.com/2024/12/10/o1-turns-pro/ "o1 Turns Pro"){.link-annotated-partial
.id-not .include-annotation}**

## `“Training Large Language Models to Reason in a Continuous Latent Space”, Hao et al 2024`{=html} {#hao-et-al-2024-section}

**[`Training Large Language Models to Reason in a Continuous Latent Space`{=html}](https://arxiv.org/abs/2412.06769#facebook "'Training Large Language Models to Reason in a Continuous Latent Space', Hao et al 2024"){.link-annotated
.id-not .include-annotation link-icon="facebook" link-icon-type="svg"
link-icon-color="#1877f2"}**

## `“Frontier Models Are Capable of In-Context Scheming”, Meinke et al 2024`{=html} {#meinke-et-al-2024-section}

**[`Frontier Models are Capable of In-Context Scheming`{=html}](https://arxiv.org/abs/2412.04984#apollo "'Frontier Models are Capable of In-Context Scheming', Meinke et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Introducing ChatGPT Pro: Broadening Usage of Frontier AI”, OpenAI 2024`{=html} {#openai-2024-section}

**[`Introducing ChatGPT Pro: Broadening usage of frontier AI`{=html}](https://openai.com/index/introducing-chatgpt-pro/ "'Introducing ChatGPT Pro: Broadening usage of frontier AI', OpenAI 2024"){.link-annotated-partial
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Frontier Models Are Capable of In-Context Scheming”, Hobbhahn et al 2024`{=html} {#hobbhahn-et-al-2024-section}

**[`Frontier Models are Capable of In-Context Scheming`{=html}](https://www.lesswrong.com/posts/8gy7c8GAPkuu6wTiX/frontier-models-are-capable-of-in-context-scheming "'Frontier Models are Capable of In-Context Scheming', Hobbhahn et al 2024"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Free Process Rewards without Process Labels”, Yuan et al 2024`{=html} {#yuan-et-al-2024-section}

**[`Free Process Rewards without Process Labels`{=html}](https://arxiv.org/abs/2412.01981 "'Free Process Rewards without Process Labels', Yuan et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models”, Ruis et al 2024`{=html} {#ruis-et-al-2024-section}

**[`Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models`{=html}](https://arxiv.org/abs/2411.12580 "'Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models', Ruis et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Mind Your Step (By Step): Chain-Of-Thought Can Reduce Performance on Tasks Where Thinking Makes Humans Worse”, Liu et al 2024`{=html} {#liu-et-al-2024-section}

**[`Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse`{=html}](https://arxiv.org/abs/2410.21333 "'Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse', Liu et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Thinking LLMs: General Instruction Following With Thought Generation”, Wu et al 2024`{=html} {#wu-et-al-2024-section}

**[`Thinking LLMs: General Instruction Following with Thought Generation`{=html}](https://arxiv.org/abs/2410.10630 "'Thinking LLMs: General Instruction Following with Thought Generation', Wu et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“When a Language Model Is Optimized for Reasoning, Does It Still Show Embers of Autoregression? An Analysis of OpenAI O1”, McCoy et al 2024`{=html} {#mccoy-et-al-2024-section}

**[`When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1`{=html}](https://arxiv.org/abs/2410.01792 "'When a language model is optimized for reasoning, does it still show embers of autoregression? An analysis of OpenAI o1', McCoy et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Evaluation of OpenAI O1: Opportunities and Challenges of AGI”, Zhong et al 2024`{=html} {#zhong-et-al-2024-1-section}

**[`Evaluation of OpenAI o1: Opportunities and Challenges of AGI`{=html}](https://arxiv.org/abs/2409.18486 "'Evaluation of OpenAI o1: Opportunities and Challenges of AGI', Zhong et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“LLMs Still Can’t Plan; Can LRMs? A Preliminary Evaluation of OpenAI’s O1 on PlanBench”, Valmeekam et al 2024`{=html} {#valmeekam-et-al-2024-section}

**[`LLMs Still Can’t Plan; Can LRMs? A Preliminary Evaluation of OpenAI’s o1 on PlanBench`{=html}](https://arxiv.org/abs/2409.13373 "'LLMs Still Can’t Plan; Can LRMs? A Preliminary Evaluation of OpenAI’s o1 on PlanBench', Valmeekam et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Training Language Models to Self-Correct via Reinforcement Learning”, Kumar et al 2024`{=html} {#kumar-et-al-2024-section}

**[`Training Language Models to Self-Correct via Reinforcement Learning`{=html}](https://arxiv.org/abs/2409.12917#deepmind "'Training Language Models to Self-Correct via Reinforcement Learning', Kumar et al 2024"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“To CoT or Not to CoT? Chain-Of-Thought Helps Mainly on Math and Symbolic Reasoning”, Sprague et al 2024`{=html} {#sprague-et-al-2024-section}

**[`To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning`{=html}](https://arxiv.org/abs/2409.12183 "'To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning', Sprague et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Critique-Out-Loud Reward Models”, Ankner et al 2024`{=html} {#ankner-et-al-2024-1-section}

**[`Critique-out-Loud Reward Models`{=html}](https://arxiv.org/abs/2408.11791#databricks "'Critique-out-Loud Reward Models', Ankner et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process”, Ye et al 2024`{=html} {#ye-et-al-2024-2-section}

**[`Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process`{=html}](https://arxiv.org/abs/2407.20311 "'Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process', Ye et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Connecting the Dots: LLMs Can Infer and Verbalize Latent Structure from Disparate Training Data”, Treutlein et al 2024`{=html} {#treutlein-et-al-2024-section}

**[`Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data`{=html}](https://arxiv.org/abs/2406.14546 "'Connecting the Dots: LLMs can Infer and Verbalize Latent Structure from Disparate Training Data', Treutlein et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?”, Lee et al 2024`{=html} {#lee-et-al-2024-1-section}

**[`Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?`{=html}](https://arxiv.org/abs/2406.13121#google "'Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?', Lee et al 2024"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“OlympicArena: Benchmarking Multi-Discipline Cognitive Reasoning for Superintelligent AI”, Huang et al 2024`{=html} {#huang-et-al-2024-3-section}

**[`OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI`{=html}](https://arxiv.org/abs/2406.12753 "'OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI', Huang et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“How Far Can Transformers Reason? The Locality Barrier and Inductive Scratchpad”, Abbe et al 2024`{=html} {#abbe-et-al-2024-section}

**[`How Far Can Transformers Reason? The Locality Barrier and Inductive Scratchpad`{=html}](https://arxiv.org/abs/2406.06467 "'How Far Can Transformers Reason? The Locality Barrier and Inductive Scratchpad', Abbe et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“OmegaPRM: Improve Mathematical Reasoning in Language Models by Automated Process Supervision”, Luo et al 2024`{=html} {#luo-et-al-2024-section}

**[`OmegaPRM: Improve Mathematical Reasoning in Language Models by Automated Process Supervision`{=html}](https://arxiv.org/abs/2406.06592#deepmind "'OmegaPRM: Improve Mathematical Reasoning in Language Models by Automated Process Supervision', Luo et al 2024"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark”, Wang et al 2024`{=html} {#wang-et-al-2024-06-section}

**[`MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark`{=html}](https://arxiv.org/abs/2406.01574 "'MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark', Wang et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“A Theoretical Understanding of Self-Correction through In-Context Alignment”, Wang et al 2024`{=html} {#wang-et-al-2024-07-section}

**[`A Theoretical Understanding of Self-Correction through In-context Alignment`{=html}](https://arxiv.org/abs/2405.18634 "'A Theoretical Understanding of Self-Correction through In-context Alignment', Wang et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Intelligent Go-Explore (IGE): Standing on the Shoulders of Giant Foundation Models”, Lu et al 2024`{=html} {#lu-et-al-2024-2-section}

**[`Intelligent Go-Explore (IGE): Standing on the Shoulders of Giant Foundation Models`{=html}](https://arxiv.org/abs/2405.15143 "'Intelligent Go-Explore (IGE): Standing on the Shoulders of Giant Foundation Models', Lu et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step”, Deng et al 2024`{=html} {#deng-et-al-2024-1-section}

**[`From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step`{=html}](https://arxiv.org/abs/2405.14838 "'From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step', Deng et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Observational Scaling Laws and the Predictability of Language Model Performance”, Ruan et al 2024`{=html} {#ruan-et-al-2024-section}

**[`Observational Scaling Laws and the Predictability of Language Model Performance`{=html}](https://arxiv.org/abs/2405.10938 "'Observational Scaling Laws and the Predictability of Language Model Performance', Ruan et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Retrieval Head Mechanistically Explains Long-Context Factuality”, Wu et al 2024`{=html} {#wu-et-al-2024-1-section}

**[`Retrieval Head Mechanistically Explains Long-Context Factuality`{=html}](https://arxiv.org/abs/2404.15574 "'Retrieval Head Mechanistically Explains Long-Context Factuality', Wu et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models”, Pfau et al 2024`{=html} {#pfau-et-al-2024-section}

**[`Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models`{=html}](https://arxiv.org/abs/2404.15758 "'Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models', Pfau et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Autonomous LLM-Driven Research from Data to Human-Verifiable Research Papers”, Ifargan et al 2024`{=html} {#ifargan-et-al-2024-section}

**[`Autonomous LLM-driven research from data to human-verifiable research papers`{=html}](https://arxiv.org/abs/2404.17605 "'Autonomous LLM-driven research from data to human-verifiable research papers', Ifargan et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Missed Connections: Lateral Thinking Puzzles for Large Language Models”, Todd et al 2024`{=html} {#todd-et-al-2024-section}

**[`Missed Connections: Lateral Thinking Puzzles for Large Language Models`{=html}](https://arxiv.org/abs/2404.11730 "'Missed Connections: Lateral Thinking Puzzles for Large Language Models', Todd et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“ChatGPT Can Predict the Future When It Tells Stories Set in the Future About the Past”, Pham & Cunningham 2024`{=html} {#pham-cunningham-2024-section}

**[`ChatGPT Can Predict the Future when it Tells Stories Set in the Future About the Past`{=html}](https://arxiv.org/abs/2404.07396 "'ChatGPT Can Predict the Future when it Tells Stories Set in the Future About the Past', Pham & Cunningham 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Visualization-Of-Thought Elicits Spatial Reasoning in Large Language Models”, Wu et al 2024`{=html} {#wu-et-al-2024-2-section}

**[`Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models`{=html}](https://arxiv.org/abs/2404.03622#microsoft "'Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models', Wu et al 2024"){.link-annotated
.id-not .include-annotation link-icon="MS"
link-icon-type="text,sans,italic"}**

## `“Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack”, Russinovich et al 2024`{=html} {#russinovich-et-al-2024-section}

**[`Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack`{=html}](https://arxiv.org/abs/2404.01833 "'Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack', Russinovich et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Do Language Models Plan Ahead for Future Tokens?”, Wu et al 2024`{=html} {#wu-et-al-2024-4-section}

**[`Do language models plan ahead for future tokens?`{=html}](https://arxiv.org/abs/2404.00859 "'Do language models plan ahead for future tokens?', Wu et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“FABLES: Evaluating Faithfulness and Content Selection in Book-Length Summarization”, Kim et al 2024`{=html} {#kim-et-al-2024-1-section}

**[`FABLES: Evaluating faithfulness and content selection in book-length summarization`{=html}](https://arxiv.org/abs/2404.01261 "'FABLES: Evaluating faithfulness and content selection in book-length summarization', Kim et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Re-Evaluating GPT-4’s Bar Exam Performance”, Martínez 2024`{=html} {#martínez-2024-section}

**[`Re-evaluating GPT-4’s bar exam performance`{=html}](https://link.springer.com/article/10.1007/s10506-024-09396-9 "'Re-evaluating GPT-4’s bar exam performance', Martínez 2024"){.link-annotated
.id-not .include-annotation link-icon="springerlink"
link-icon-type="svg"}**

## `“Long-Form Factuality in Large Language Models”, Wei et al 2024`{=html} {#wei-et-al-2024-1-section}

**[`Long-form factuality in large language models`{=html}](https://arxiv.org/abs/2403.18802#deepmind "'Long-form factuality in large language models', Wei et al 2024"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Don’t Trust: Verify—Grounding LLM Quantitative Reasoning With Autoformalization”, Zhou et al 2024`{=html} {#zhou-et-al-2024-section}

**[`Don’t Trust: Verify—Grounding LLM Quantitative Reasoning with Autoformalization`{=html}](https://arxiv.org/abs/2403.18120#google "'Don’t Trust: Verify—Grounding LLM Quantitative Reasoning with Autoformalization', Zhou et al 2024"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking”, Zelikman et al 2024`{=html} {#zelikman-et-al-2024-section}

**[`Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking`{=html}](https://arxiv.org/abs/2403.09629 "'Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking', Zelikman et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“RNNs Are Not Transformers (Yet): The Key Bottleneck on In-Context Retrieval”, Wen et al 2024`{=html} {#wen-et-al-2024-2-section}

**[`RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval`{=html}](https://arxiv.org/abs/2402.18510 "'RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval', Wen et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Tokenization Counts: the Impact of Tokenization on Arithmetic in Frontier LLMs”, Singh & Strouse 2024`{=html} {#singh-strouse-2024-section}

**[`Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs`{=html}](https://arxiv.org/abs/2402.14903 "'Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs', Singh & Strouse 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Chain-Of-Thought Empowers Transformers to Solve Inherently Serial Problems”, Li et al 2024`{=html} {#li-et-al-2024-09-section}

**[`Chain-of-Thought Empowers Transformers to Solve Inherently Serial Problems`{=html}](https://arxiv.org/abs/2402.12875 "'Chain-of-Thought Empowers Transformers to Solve Inherently Serial Problems', Li et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models”, Levy et al 2024`{=html} {#levy-et-al-2024-section}

**[`Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models`{=html}](https://arxiv.org/abs/2402.14848 "'Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models', Levy et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Why Are Sensitive Functions Hard for Transformers?”, Hahn & Rofin 2024`{=html} {#hahn-rofin-2024-section}

**[`Why are Sensitive Functions Hard for Transformers?`{=html}](https://arxiv.org/abs/2402.09963 "'Why are Sensitive Functions Hard for Transformers?', Hahn & Rofin 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Chain-Of-Thought Reasoning Without Prompting”, Wang & Zhou 2024`{=html} {#wang-zhou-2024-section}

**[`Chain-of-Thought Reasoning Without Prompting`{=html}](https://arxiv.org/abs/2402.10200#deepmind "'Chain-of-Thought Reasoning Without Prompting', Wang & Zhou 2024"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“V-STaR: Training Verifiers for Self-Taught Reasoners”, Hosseini et al 2024`{=html} {#hosseini-et-al-2024-section}

**[`V-STaR: Training Verifiers for Self-Taught Reasoners`{=html}](https://arxiv.org/abs/2402.06457 "'V-STaR: Training Verifiers for Self-Taught Reasoners', Hosseini et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“More Agents Is All You Need”, Li et al 2024`{=html} {#li-et-al-2024-10-section}

**[`More Agents Is All You Need`{=html}](https://arxiv.org/abs/2402.05120#tencent "'More Agents Is All You Need', Li et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Knowledge Distillation of Black-Box Large Language Models”, Chen et al 2024`{=html} {#chen-et-al-2024-section}

**[`Knowledge Distillation of Black-Box Large Language Models`{=html}](https://arxiv.org/abs/2401.07013 "'Knowledge Distillation of Black-Box Large Language Models', Chen et al 2024"){.link-modified-recently
.link-annotated .id-not .include-annotation link-icon="𝛘"
link-icon-type="text" link-icon-color="#b31b1b"}**

## `“The Impact of Reasoning Step Length on Large Language Models”, Jin et al 2024`{=html} {#jin-et-al-2024-4-section}

**[`The Impact of Reasoning Step Length on Large Language Models`{=html}](https://arxiv.org/abs/2401.04925 "'The Impact of Reasoning Step Length on Large Language Models', Jin et al 2024"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models Play <em>StarCraft II</em>: Benchmarks and A Chain of Summarization Approach”, Ma et al 2023`{=html} {#ma-et-al-2023-section}

**[`Large Language Models Play <em>StarCraft II</em>: Benchmarks and A Chain of Summarization Approach`{=html}](https://arxiv.org/abs/2312.11865 "'Large Language Models Play <em>StarCraft II</em>: Benchmarks and A Chain of Summarization Approach', Ma et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Math-Shepherd: Verify and Reinforce LLMs Step-By-Step without Human Annotations”, Wang et al 2023`{=html} {#wang-et-al-2023-section}

**[`Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations`{=html}](https://arxiv.org/abs/2312.08935 "'Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations', Wang et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Beyond Human Data: Scaling Self-Training for Problem-Solving With Language Models (ReST<sup>EM</sup>)”, Singh et al 2023`{=html} {#singh-et-al-2023-3-section}

**[`Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models (ReST<sup>EM</sup>)`{=html}](https://arxiv.org/abs/2312.06585#deepmind "'Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models (ReST<sup>EM</sup>)', Singh et al 2023"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Tree of Attacks (TAP): Jailbreaking Black-Box LLMs Automatically”, Mehrotra et al 2023`{=html} {#mehrotra-et-al-2023-section}

**[`Tree of Attacks (TAP): Jailbreaking Black-Box LLMs Automatically`{=html}](https://arxiv.org/abs/2312.02119 "'Tree of Attacks (TAP): Jailbreaking Black-Box LLMs Automatically', Mehrotra et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Universal Self-Consistency for Large Language Model Generation”, Chen et al 2023`{=html} {#chen-et-al-2023-03-section}

**[`Universal Self-Consistency for Large Language Model Generation`{=html}](https://arxiv.org/abs/2311.17311#deepmind "'Universal Self-Consistency for Large Language Model Generation', Chen et al 2023"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine”, Nori et al 2023`{=html} {#nori-et-al-2023-section}

**[`Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine`{=html}](https://arxiv.org/abs/2311.16452#microsoft "'Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine', Nori et al 2023"){.link-annotated
.id-not .include-annotation link-icon="MS"
link-icon-type="text,sans,italic"}**

## `“Training Chain-Of-Thought via Latent-Variable Inference”, Phan et al 2023`{=html} {#phan-et-al-2023-section}

**[`Training Chain-of-Thought via Latent-Variable Inference`{=html}](https://arxiv.org/abs/2312.02179 "'Training Chain-of-Thought via Latent-Variable Inference', Phan et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks”, Ramesh et al 2023`{=html} {#ramesh-et-al-2023-section}

**[`Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks`{=html}](https://arxiv.org/abs/2311.12997 "'Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks', Ramesh et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“On Measuring Faithfulness or Self-Consistency of Natural Language Explanations”, Parcalabescu & Frank 2023`{=html} {#parcalabescu-frank-2023-section}

**[`On Measuring Faithfulness or Self-consistency of Natural Language Explanations`{=html}](https://arxiv.org/abs/2311.07466 "'On Measuring Faithfulness or Self-consistency of Natural Language Explanations', Parcalabescu & Frank 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Zero-Shot Goal-Directed Dialogue via RL on Imagined Conversations”, Hong et al 2023`{=html} {#hong-et-al-2023-section}

**[`Zero-Shot Goal-Directed Dialogue via RL on Imagined Conversations`{=html}](https://arxiv.org/abs/2311.05584 "'Zero-Shot Goal-Directed Dialogue via RL on Imagined Conversations', Hong et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models Can Strategically Deceive Their Users When Put Under Pressure”, Scheurer et al 2023`{=html} {#scheurer-et-al-2023-section}

**[`Large Language Models can Strategically Deceive their Users when Put Under Pressure`{=html}](https://arxiv.org/abs/2311.07590#apollo "'Large Language Models can Strategically Deceive their Users when Put Under Pressure', Scheurer et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves”, Deng et al 2023`{=html} {#deng-et-al-2023-1-section}

**[`Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves`{=html}](https://arxiv.org/abs/2311.04205 "'Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves', Deng et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Everything of Thoughts: Defying the Law of Penrose Triangle for Thought Generation”, Ding et al 2023`{=html} {#ding-et-al-2023-2-section}

**[`Everything of Thoughts: Defying the Law of Penrose Triangle for Thought Generation`{=html}](https://arxiv.org/abs/2311.04254 "'Everything of Thoughts: Defying the Law of Penrose Triangle for Thought Generation', Ding et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Qwen3: Think Deeper, Act Faster”, Alibaba 2023`{=html} {#alibaba-2023-section}

**[`Qwen3: Think Deeper, Act Faster`{=html}](https://qwenlm.github.io/blog/qwen3/#alibaba "'Qwen3: Think Deeper, Act Faster', Alibaba 2023"){.id-not
.include-annotation}**

::: aux-links-transclude-file
::: collapse
**View External Link**:

[`https://qwenlm.github.io/blog/qwen3/#alibaba`](https://qwenlm.github.io/blog/qwen3/#alibaba){.id-not
.link-annotated-not .include-content .include-lazy}
:::
:::

## `“Implicit Chain-Of-Thought Reasoning via Knowledge Distillation”, Deng et al 2023`{=html} {#deng-et-al-2023-2-section}

**[`Implicit Chain-of-Thought Reasoning via Knowledge Distillation`{=html}](https://arxiv.org/abs/2311.01460 "'Implicit Chain-of-Thought Reasoning via Knowledge Distillation', Deng et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Preventing Language Models From Hiding Their Reasoning”, Roger & Greenblatt 2023`{=html} {#roger-greenblatt-2023-section}

**[`Preventing Language Models From Hiding Their Reasoning`{=html}](https://arxiv.org/abs/2310.18512#redwood "'Preventing Language Models From Hiding Their Reasoning', Roger & Greenblatt 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Branch-Solve-Merge Improves Large Language Model Evaluation and Generation”, Saha et al 2023`{=html} {#saha-et-al-2023-section}

**[`Branch-Solve-Merge Improves Large Language Model Evaluation and Generation`{=html}](https://arxiv.org/abs/2310.15123 "'Branch-Solve-Merge Improves Large Language Model Evaluation and Generation', Saha et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Can GPT Models Be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on Mock CFA Exams”, Callanan et al 2023`{=html} {#callanan-et-al-2023-section}

**[`Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams`{=html}](https://arxiv.org/abs/2310.08678 "'Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams', Callanan et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“The Expressive Power of Transformers With Chain-Of-Thought”, Merrill & Sabharwal 2023`{=html} {#merrill-sabharwal-2023-section}

**[`The Expressive Power of Transformers with Chain-of-Thought`{=html}](https://arxiv.org/abs/2310.07923 "'The Expressive Power of Transformers with Chain-of-Thought', Merrill & Sabharwal 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”, Zhou et al 2023`{=html} {#zhou-et-al-2023-04-section}

**[`Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models`{=html}](https://arxiv.org/abs/2310.04406 "'Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models', Zhou et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models Cannot Self-Correct Reasoning Yet”, Huang et al 2023`{=html} {#huang-et-al-2023-2-section}

**[`Large Language Models Cannot Self-Correct Reasoning Yet`{=html}](https://arxiv.org/abs/2310.01798 "'Large Language Models Cannot Self-Correct Reasoning Yet', Huang et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Think Before You Speak: Training Language Models With Pause Tokens”, Goyal et al 2023`{=html} {#goyal-et-al-2023-section}

**[`Think before you speak: Training Language Models With Pause Tokens`{=html}](https://arxiv.org/abs/2310.02226 "'Think before you speak: Training Language Models With Pause Tokens', Goyal et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Embers of Autoregression: Understanding Large Language Models Through the Problem They Are Trained to Solve”, McCoy et al 2023`{=html} {#mccoy-et-al-2023-section}

**[`Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve`{=html}](https://arxiv.org/abs/2309.13638 "'Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve', McCoy et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Contrastive Decoding Improves Reasoning in Large Language Models”, O’Brien & Lewis 2023`{=html} {#obrien-lewis-2023-section}

**[`Contrastive Decoding Improves Reasoning in Large Language Models`{=html}](https://arxiv.org/abs/2309.09117#facebook "'Contrastive Decoding Improves Reasoning in Large Language Models', O’Brien & Lewis 2023"){.link-annotated
.id-not .include-annotation link-icon="facebook" link-icon-type="svg"
link-icon-color="#1877f2"}**

## `“Re2: Re-Reading Improves Reasoning in Large Language Models”, Xu et al 2023`{=html} {#xu-et-al-2023-section}

**[`Re2: Re-Reading Improves Reasoning in Large Language Models`{=html}](https://arxiv.org/abs/2309.06275 "'Re2: Re-Reading Improves Reasoning in Large Language Models', Xu et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“From Sparse to Dense: GPT-4 Summarization With Chain of Density (CoD) Prompting”, Adams et al 2023`{=html} {#adams-et-al-2023-section}

**[`From Sparse to Dense: GPT-4 Summarization with Chain of Density (CoD) Prompting`{=html}](https://arxiv.org/abs/2309.04269 "'From Sparse to Dense: GPT-4 Summarization with Chain of Density (CoD) Prompting', Adams et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Graph of Thoughts: Solving Elaborate Problems With Large Language Models”, Besta et al 2023`{=html} {#besta-et-al-2023-section}

**[`Graph of Thoughts: Solving Elaborate Problems with Large Language Models`{=html}](https://arxiv.org/abs/2308.09687 "'Graph of Thoughts: Solving Elaborate Problems with Large Language Models', Besta et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Solving Challenging Math Word Problems Using GPT-4 Code Interpreter With Code-Based Self-Verification”, Zhou et al 2023`{=html} {#zhou-et-al-2023-06-section}

**[`Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification`{=html}](https://arxiv.org/abs/2308.07921 "'Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification', Zhou et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Scaling Relationship on Learning Mathematical Reasoning With Large Language Models”, Yuan et al 2023`{=html} {#yuan-et-al-2023-section}

**[`Scaling Relationship on Learning Mathematical Reasoning with Large Language Models`{=html}](https://arxiv.org/abs/2308.01825#alibaba "'Scaling Relationship on Learning Mathematical Reasoning with Large Language Models', Yuan et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Android in the Wild: A Large-Scale Dataset for Android Device Control”, Rawles et al 2023`{=html} {#rawles-et-al-2023-section}

**[`Android in the Wild: A Large-Scale Dataset for Android Device Control`{=html}](https://arxiv.org/abs/2307.10088#google "'Android in the Wild: A Large-Scale Dataset for Android Device Control', Rawles et al 2023"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“LLMs As Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines With LLMs”, Wu et al 2023`{=html} {#wu-et-al-2023-2-section}

**[`LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs`{=html}](https://arxiv.org/abs/2307.10168 "'LLMs as Workers in Human-Computational Algorithms? Replicating Crowdsourcing Pipelines with LLMs', Wu et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“TableGPT: Towards Unifying Tables, Nature Language and Commands into One GPT”, Zha et al 2023`{=html} {#zha-et-al-2023-section}

**[`TableGPT: Towards Unifying Tables, Nature Language and Commands into One GPT`{=html}](https://arxiv.org/abs/2307.08674 "'TableGPT: Towards Unifying Tables, Nature Language and Commands into One GPT', Zha et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Question Decomposition Improves the Faithfulness of Model-Generated Reasoning”, Radhakrishnan et al 2023`{=html} {#radhakrishnan-et-al-2023-section}

**[`Question Decomposition Improves the Faithfulness of Model-Generated Reasoning`{=html}](https://arxiv.org/abs/2307.11768#anthropic "'Question Decomposition Improves the Faithfulness of Model-Generated Reasoning', Radhakrishnan et al 2023"){.link-annotated
.id-not .include-annotation link-icon="anthropic" link-icon-type="svg"
link-icon-color="#d4a27f"}**

## `“Measuring Faithfulness in Chain-Of-Thought Reasoning”, Lanham et al 2023`{=html} {#lanham-et-al-2023-section}

**[`Measuring Faithfulness in Chain-of-Thought Reasoning`{=html}](https://arxiv.org/abs/2307.13702 "'Measuring Faithfulness in Chain-of-Thought Reasoning', Lanham et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration”, Wang et al 2023`{=html} {#wang-et-al-2023-12-section}

**[`Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration`{=html}](https://arxiv.org/abs/2307.05300#microsoft "'Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration', Wang et al 2023"){.link-annotated
.id-not .include-annotation link-icon="MS"
link-icon-type="text,sans,italic"}**

## `“Explaining Competitive-Level Programming Solutions Using LLMs”, Li et al 2023`{=html} {#li-et-al-2023-05-section}

**[`Explaining Competitive-Level Programming Solutions using LLMs`{=html}](https://arxiv.org/abs/2307.05337 "'Explaining Competitive-Level Programming Solutions using LLMs', Li et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Teaching Arithmetic to Small Transformers”, Lee et al 2023`{=html} {#lee-et-al-2023-2-section}

**[`Teaching Arithmetic to Small Transformers`{=html}](https://arxiv.org/abs/2307.03381 "'Teaching Arithmetic to Small Transformers', Lee et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Language Models Are Weak Learners”, Manikandan et al 2023`{=html} {#manikandan-et-al-2023-section}

**[`Language models are weak learners`{=html}](https://arxiv.org/abs/2306.14101 "'Language models are weak learners', Manikandan et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Let’s Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning”, Ma et al 2023`{=html} {#ma-et-al-2023-2-section}

**[`Let’s Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning`{=html}](https://arxiv.org/abs/2306.14308#google "'Let’s Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning', Ma et al 2023"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“GKD: Generalized Knowledge Distillation for Auto-Regressive Sequence Models”, Agarwal et al 2023`{=html} {#agarwal-et-al-2023-2-section}

**[`GKD: Generalized Knowledge Distillation for Auto-regressive Sequence Models`{=html}](https://arxiv.org/abs/2306.13649#deepmind "'GKD: Generalized Knowledge Distillation for Auto-regressive Sequence Models', Agarwal et al 2023"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought”, Wong et al 2023`{=html} {#wong-et-al-2023-2-section}

**[`From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought`{=html}](https://arxiv.org/abs/2306.12672 "'From Word Models to World Models: Translating from Natural Language to the Probabilistic Language of Thought', Wong et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models As Tax Attorneys: A Case Study in Legal Capabilities Emergence”, Nay et al 2023`{=html} {#nay-et-al-2023-section}

**[`Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence`{=html}](https://arxiv.org/abs/2306.07075 "'Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence', Nay et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Iterative Translation Refinement With Large Language Models”, Chen et al 2023`{=html} {#chen-et-al-2023-11-section}

**[`Iterative Translation Refinement with Large Language Models`{=html}](https://arxiv.org/abs/2306.03856 "'Iterative Translation Refinement with Large Language Models', Chen et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Thought Cloning: Learning to Think While Acting by Imitating Human Thinking”, Hu & Clune 2023`{=html} {#hu-clune-2023-section}

**[`Thought Cloning: Learning to Think while Acting by Imitating Human Thinking`{=html}](https://arxiv.org/abs/2306.00323 "'Thought Cloning: Learning to Think while Acting by Imitating Human Thinking', Hu & Clune 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Let’s Verify Step by Step”, Lightman et al 2023`{=html} {#lightman-et-al-2023-section}

**[`Let’s Verify Step by Step`{=html}](https://arxiv.org/abs/2305.20050#openai "'Let’s Verify Step by Step', Lightman et al 2023"){.link-annotated
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Towards Revealing the Mystery behind Chain-Of-Thought: A Theoretical Perspective”, Feng et al 2023`{=html} {#feng-et-al-2023-2-section}

**[`Towards Revealing the Mystery behind Chain-of-Thought: A Theoretical Perspective`{=html}](https://arxiv.org/abs/2305.15408 "'Towards Revealing the Mystery behind Chain-of-Thought: A Theoretical Perspective', Feng et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Improving Factuality and Reasoning in Language Models through Multiagent Debate”, Du et al 2023`{=html} {#du-et-al-2023-3-section}

**[`Improving Factuality and Reasoning in Language Models through Multiagent Debate`{=html}](https://arxiv.org/abs/2305.14325 "'Improving Factuality and Reasoning in Language Models through Multiagent Debate', Du et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“How Language Model Hallucinations Can Snowball”, Zhang et al 2023`{=html} {#zhang-et-al-2023-14-section}

**[`How Language Model Hallucinations Can Snowball`{=html}](https://arxiv.org/abs/2305.13534 "'How Language Model Hallucinations Can Snowball', Zhang et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Tree of Thoughts (ToT): Deliberate Problem Solving With Large Language Models”, Yao et al 2023`{=html} {#yao-et-al-2023-section}

**[`Tree of Thoughts (ToT): Deliberate Problem Solving with Large Language Models`{=html}](https://arxiv.org/abs/2305.10601#deepmind "'Tree of Thoughts (ToT): Deliberate Problem Solving with Large Language Models', Yao et al 2023"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Large Language Model Programs”, Schlag et al 2023`{=html} {#schlag-et-al-2023-section}

**[`Large Language Model Programs`{=html}](https://arxiv.org/abs/2305.05364#facebook "'Large Language Model Programs', Schlag et al 2023"){.link-annotated
.id-not .include-annotation link-icon="facebook" link-icon-type="svg"
link-icon-color="#1877f2"}**

## `“Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-Of-Thought Prompting”, Turpin et al 2023`{=html} {#turpin-et-al-2023-section}

**[`Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting`{=html}](https://arxiv.org/abs/2305.04388 "'Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting', Turpin et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Distilling Step-By-Step! Outperforming Larger Language Models With Less Training Data and Smaller Model Sizes”, Hsieh et al 2023`{=html} {#hsieh-et-al-2023-2-section}

**[`Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes`{=html}](https://arxiv.org/abs/2305.02301#google "'Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes', Hsieh et al 2023"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Decomposition Enhances Reasoning via Self-Evaluation Guided Decoding”, Xie et al 2023`{=html} {#xie-et-al-2023-2-section}

**[`Decomposition Enhances Reasoning via Self-Evaluation Guided Decoding`{=html}](https://arxiv.org/abs/2305.00633 "'Decomposition Enhances Reasoning via Self-Evaluation Guided Decoding', Xie et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“LLM+P: Empowering Large Language Models With Optimal Planning Proficiency”, Liu et al 2023`{=html} {#liu-et-al-2023-16-section}

**[`LLM+P: Empowering Large Language Models with Optimal Planning Proficiency`{=html}](https://arxiv.org/abs/2304.11477 "'LLM+P: Empowering Large Language Models with Optimal Planning Proficiency', Liu et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Boosting Theory-Of-Mind Performance in Large Language Models via Prompting”, Moghaddam & Honey 2023`{=html} {#moghaddam-honey-2023-section}

**[`Boosting Theory-of-Mind Performance in Large Language Models via Prompting`{=html}](https://arxiv.org/abs/2304.11490 "'Boosting Theory-of-Mind Performance in Large Language Models via Prompting', Moghaddam & Honey 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Think Before You Act: Unified Policy for Interleaving Language Reasoning With Actions”, Mezghani et al 2023`{=html} {#mezghani-et-al-2023-section}

**[`Think Before You Act: Unified Policy for Interleaving Language Reasoning with Actions`{=html}](https://arxiv.org/abs/2304.11063#facebook "'Think Before You Act: Unified Policy for Interleaving Language Reasoning with Actions', Mezghani et al 2023"){.link-annotated
.id-not .include-annotation link-icon="facebook" link-icon-type="svg"
link-icon-color="#1877f2"}**

## `“Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games”`{=html} {#_hKTsaimv-section}

**[`Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games`{=html}](https://www.lesswrong.com/posts/M6dXdCbdoLSpHt8v3/corrupted-by-reasoning-reasoning-language-models-become-free "Corrupted by Reasoning: Reasoning Language Models Become Free-Riders in Public Goods Games"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Language Models Can Solve Computer Tasks”, Kim et al 2023`{=html} {#kim-et-al-2023-6-section}

**[`Language Models can Solve Computer Tasks`{=html}](https://arxiv.org/abs/2303.17491 "'Language Models can Solve Computer Tasks', Kim et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Reflexion: Language Agents With Verbal Reinforcement Learning”, Shinn et al 2023`{=html} {#shinn-et-al-2023-section}

**[`Reflexion: Language Agents with Verbal Reinforcement Learning`{=html}](https://arxiv.org/abs/2303.11366 "'Reflexion: Language Agents with Verbal Reinforcement Learning', Shinn et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“How Well Do Large Language Models Perform in Arithmetic Tasks?”, Yuan et al 2023`{=html} {#yuan-et-al-2023-2-section}

**[`How well do Large Language Models perform in Arithmetic tasks?`{=html}](https://arxiv.org/abs/2304.02015#alibaba "'How well do Large Language Models perform in Arithmetic tasks?', Yuan et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models”, Manakul et al 2023`{=html} {#manakul-et-al-2023-section}

**[`SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models`{=html}](https://arxiv.org/abs/2303.08896 "'SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models', Manakul et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Language Is Not All You Need: Aligning Perception With Language Models (Kosmos-1)”, Huang et al 2023`{=html} {#huang-et-al-2023-7-section}

**[`Language Is Not All You Need: Aligning Perception with Language Models (Kosmos-1)`{=html}](https://arxiv.org/abs/2302.14045#microsoft "'Language Is Not All You Need: Aligning Perception with Language Models (Kosmos-1)', Huang et al 2023"){.link-annotated
.id-not .include-annotation link-icon="MS"
link-icon-type="text,sans,italic"}**

## `“Multimodal Chain-Of-Thought Reasoning in Language Models”, Zhang et al 2023`{=html} {#zhang-et-al-2023-19-section}

**[`Multimodal Chain-of-Thought Reasoning in Language Models`{=html}](https://arxiv.org/abs/2302.00923#amazon "'Multimodal Chain-of-Thought Reasoning in Language Models', Zhang et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Faithful Chain-Of-Thought Reasoning”, Lyu et al 2023`{=html} {#lyu-et-al-2023-2-section}

**[`Faithful Chain-of-Thought Reasoning`{=html}](https://arxiv.org/abs/2301.13379 "'Faithful Chain-of-Thought Reasoning', Lyu et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models Are Versatile Decomposers: Decompose Evidence and Questions for Table-Based Reasoning”, Ye et al 2023`{=html} {#ye-et-al-2023-section}

**[`Large Language Models are Versatile Decomposers: Decompose Evidence and Questions for Table-based Reasoning`{=html}](https://arxiv.org/abs/2301.13808#alibaba "'Large Language Models are Versatile Decomposers: Decompose Evidence and Questions for Table-based Reasoning', Ye et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“ChatGPT Goes to Law School”, Choi et al 2023`{=html} {#choi-et-al-2023-section}

**[`ChatGPT Goes to Law School`{=html}](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335905 "'ChatGPT Goes to Law School', Choi et al 2023"){.link-annotated
.id-not .include-annotation link-icon="SSRN" link-icon-type="text,quad"
link-icon-color="#007398"}**

## `“Large Language Models As Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards”, Nay 2023`{=html} {#nay-2023-section}

**[`Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards`{=html}](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335945 "'Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards', Nay 2023"){.link-annotated
.id-not .include-annotation link-icon="SSRN" link-icon-type="text,quad"
link-icon-color="#007398"}**

## `“Interactive-Chain-Prompting (INTERCPT): Ambiguity Resolution for Crosslingual Conditional Generation With Interaction”, Pilault et al 2023`{=html} {#pilault-et-al-2023-section}

**[`Interactive-Chain-Prompting (INTERCPT): Ambiguity Resolution for Crosslingual Conditional Generation with Interaction`{=html}](https://arxiv.org/abs/2301.10309#google "'Interactive-Chain-Prompting (INTERCPT): Ambiguity Resolution for Crosslingual Conditional Generation with Interaction', Pilault et al 2023"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Iterated Decomposition: Improving Science Q&amp;A by Supervising Reasoning Processes”, Reppert et al 2023`{=html} {#reppert-et-al-2023-section}

**[`Iterated Decomposition: Improving Science Q&amp;A by Supervising Reasoning Processes`{=html}](https://arxiv.org/abs/2301.01751#elicit "'Iterated Decomposition: Improving Science Q&amp;A by Supervising Reasoning Processes', Reppert et al 2023"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Solving Math Word Problems With Process & Outcome-Based Feedback”, Uesato et al 2022`{=html} {#uesato-et-al-2022-section}

**[`Solving math word problems with process & outcome-based feedback`{=html}](https://arxiv.org/abs/2211.14275#deepmind "'Solving math word problems with process & outcome-based feedback', Uesato et al 2022"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“PAL: Program-Aided Language Models”, Gao et al 2022`{=html} {#gao-et-al-2022-4-section}

**[`PAL: Program-aided Language Models`{=html}](https://arxiv.org/abs/2211.10435 "'PAL: Program-aided Language Models', Gao et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Measuring Progress on Scalable Oversight for Large Language Models”, Bowman et al 2022`{=html} {#bowman-et-al-2022-section}

**[`Measuring Progress on Scalable Oversight for Large Language Models`{=html}](https://arxiv.org/abs/2211.03540#anthropic "'Measuring Progress on Scalable Oversight for Large Language Models', Bowman et al 2022"){.link-annotated
.id-not .include-annotation link-icon="anthropic" link-icon-type="svg"
link-icon-color="#d4a27f"}**

## `“U-PaLM: Transcending Scaling Laws With 0.1% Extra Compute”, Tay et al 2022`{=html} {#tay-et-al-2022-upalm-section}

**[`U-PaLM: Transcending Scaling Laws with 0.1% Extra Compute`{=html}](https://arxiv.org/abs/2210.11399#google "'U-PaLM: Transcending Scaling Laws with 0.1% Extra Compute', Tay et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Large Language Models Can Self-Improve”, Huang et al 2022`{=html} {#huang-et-al-2022-2-section}

**[`Large Language Models Can Self-Improve`{=html}](https://arxiv.org/abs/2210.11610#google "'Large Language Models Can Self-Improve', Huang et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Challenging BIG-Bench Tasks (BBH) and Whether Chain-Of-Thought Can Solve Them”, Suzgun et al 2022`{=html} {#suzgun-et-al-2022-1-section}

**[`Challenging BIG-Bench Tasks (BBH) and Whether Chain-of-Thought Can Solve Them`{=html}](https://arxiv.org/abs/2210.09261#google "'Challenging BIG-Bench Tasks (BBH) and Whether Chain-of-Thought Can Solve Them', Suzgun et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Self-Ask: Measuring and Narrowing the Compositionality Gap in Language Models (Bamboogle)”, Press et al 2022`{=html} {#press-et-al-2022-section}

**[`Self-Ask: Measuring and Narrowing the Compositionality Gap in Language Models (Bamboogle)`{=html}](https://arxiv.org/abs/2210.03350#allen "'Self-Ask: Measuring and Narrowing the Compositionality Gap in Language Models (Bamboogle)', Press et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Language Models Are Multilingual Chain-Of-Thought Reasoners”, Shi et al 2022`{=html} {#shi-et-al-2022-2-section}

**[`Language Models are Multilingual Chain-of-Thought Reasoners`{=html}](https://arxiv.org/abs/2210.03057#google "'Language Models are Multilingual Chain-of-Thought Reasoners', Shi et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“ReAct: Synergizing Reasoning and Acting in Language Models”, Yao et al 2022`{=html} {#yao-et-al-2022-1-section}

**[`ReAct: Synergizing Reasoning and Acting in Language Models`{=html}](https://arxiv.org/abs/2210.03629#google "'ReAct: Synergizing Reasoning and Acting in Language Models', Yao et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Context Distillation: Learning by Distilling Context”, Snell et al 2022`{=html} {#snell-et-al-2022-section}

**[`Context Distillation: Learning by Distilling Context`{=html}](https://arxiv.org/abs/2209.15189 "'Context Distillation: Learning by Distilling Context', Snell et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Dynamic Prompt Learning via Policy Gradient for Semi-Structured Mathematical Reasoning”, Lu et al 2022`{=html} {#lu-et-al-2022-3-section}

**[`Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning`{=html}](https://arxiv.org/abs/2209.14610 "'Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning', Lu et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“FOLIO: Natural Language Reasoning With First-Order Logic”, Han et al 2022`{=html} {#han-et-al-2022-section}

**[`FOLIO: Natural Language Reasoning with First-Order Logic`{=html}](https://arxiv.org/abs/2209.00840 "'FOLIO: Natural Language Reasoning with First-Order Logic', Han et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Faithful Reasoning Using Large Language Models”, Creswell & Shanahan 2022`{=html} {#creswell-shanahan-2022-section}

**[`Faithful Reasoning Using Large Language Models`{=html}](https://arxiv.org/abs/2208.14271#deepmind "'Faithful Reasoning Using Large Language Models', Creswell & Shanahan 2022"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Limitations of Language Models in Arithmetic and Symbolic Induction”, Qian et al 2022`{=html} {#qian-et-al-2022-1-section}

**[`Limitations of Language Models in Arithmetic and Symbolic Induction`{=html}](https://arxiv.org/abs/2208.05051 "'Limitations of Language Models in Arithmetic and Symbolic Induction', Qian et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Language Models Can Teach Themselves to Program Better”, Haluptzok et al 2022`{=html} {#haluptzok-et-al-2022-section}

**[`Language Models Can Teach Themselves to Program Better`{=html}](https://arxiv.org/abs/2207.14502#microsoft "'Language Models Can Teach Themselves to Program Better', Haluptzok et al 2022"){.link-annotated
.id-not .include-annotation link-icon="MS"
link-icon-type="text,sans,italic"}**

## `“Language Model Cascades”, Dohan et al 2022`{=html} {#dohan-et-al-2022-section}

**[`Language Model Cascades`{=html}](https://arxiv.org/abs/2207.10342#google "'Language Model Cascades', Dohan et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“CodeT: Code Generation With Generated Tests”, Chen et al 2022`{=html} {#chen-et-al-2022-codet-section}

**[`CodeT: Code Generation with Generated Tests`{=html}](https://arxiv.org/abs/2207.10397#microsoft "'CodeT: Code Generation with Generated Tests', Chen et al 2022"){.link-annotated
.id-not .include-annotation link-icon="MS"
link-icon-type="text,sans,italic"}**

## `“Can Large Language Models Reason about Medical Questions?”, Liévin et al 2022`{=html} {#liévin-et-al-2022-section}

**[`Can large language models reason about medical questions?`{=html}](https://arxiv.org/abs/2207.08143 "'Can large language models reason about medical questions?', Liévin et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Inner Monologue: Embodied Reasoning through Planning With Language Models”, Huang et al 2022`{=html} {#huang-et-al-2022-5-section}

**[`Inner Monologue: Embodied Reasoning through Planning with Language Models`{=html}](https://arxiv.org/abs/2207.05608#google "'Inner Monologue: Embodied Reasoning through Planning with Language Models', Huang et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Exploring Length Generalization in Large Language Models”, Anil et al 2022`{=html} {#anil-et-al-2022-section}

**[`Exploring Length Generalization in Large Language Models`{=html}](https://arxiv.org/abs/2207.04901#google "'Exploring Length Generalization in Large Language Models', Anil et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Language Models (Mostly) Know What They Know”, Kadavath et al 2022`{=html} {#kadavath-et-al-2022-section}

**[`Language Models (Mostly) Know What They Know`{=html}](https://arxiv.org/abs/2207.05221#anthropic "'Language Models (Mostly) Know What They Know', Kadavath et al 2022"){.link-annotated
.id-not .include-annotation link-icon="anthropic" link-icon-type="svg"
link-icon-color="#d4a27f"}**

## `“Neural Networks and the Chomsky Hierarchy”, Delétang et al 2022`{=html} {#delétang-et-al-2022-section}

**[`Neural Networks and the Chomsky Hierarchy`{=html}](https://arxiv.org/abs/2207.02098#deepmind "'Neural Networks and the Chomsky Hierarchy', Delétang et al 2022"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Solving Quantitative Reasoning Problems With Language Models”, Lewkowycz et al 2022`{=html} {#lewkowycz-et-al-2022-section}

**[`Solving Quantitative Reasoning Problems with Language Models`{=html}](https://arxiv.org/abs/2206.14858#google "'Solving Quantitative Reasoning Problems with Language Models', Lewkowycz et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Maieutic Prompting: Logically Consistent Reasoning With Recursive Explanations”, Jung et al 2022`{=html} {#jung-et-al-2022-section}

**[`Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations`{=html}](https://arxiv.org/abs/2205.11822#allen "'Maieutic Prompting: Logically Consistent Reasoning with Recursive Explanations', Jung et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Large Language Models Are Zero-Shot Reasoners”, Kojima et al 2022`{=html} {#kojima-et-al-2022-section}

**[`Large Language Models are Zero-Shot Reasoners`{=html}](https://arxiv.org/abs/2205.11916 "'Large Language Models are Zero-Shot Reasoners', Kojima et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Instruction Induction: From Few Examples to Natural Language Task Descriptions”, Honovich et al 2022`{=html} {#honovich-et-al-2022-2-section}

**[`Instruction Induction: From Few Examples to Natural Language Task Descriptions`{=html}](https://arxiv.org/abs/2205.10782 "'Instruction Induction: From Few Examples to Natural Language Task Descriptions', Honovich et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Least-To-Most Prompting Enables Complex Reasoning in Large Language Models”, Zhou et al 2022`{=html} {#zhou-et-al-2022-1-section}

**[`Least-to-Most Prompting Enables Complex Reasoning in Large Language Models`{=html}](https://arxiv.org/abs/2205.10625#google "'Least-to-Most Prompting Enables Complex Reasoning in Large Language Models', Zhou et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Dialog Inpainting: Turning Documents into Dialogues”, Dai et al 2022`{=html} {#dai-et-al-2022-2-section}

**[`Dialog Inpainting: Turning Documents into Dialogues`{=html}](https://arxiv.org/abs/2205.09073#google "'Dialog Inpainting: Turning Documents into Dialogues', Dai et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“UL2: Unifying Language Learning Paradigms”, Tay et al 2022`{=html} {#tay-et-al-2022-ul2-section}

**[`UL2: Unifying Language Learning Paradigms`{=html}](https://arxiv.org/abs/2205.05131#google "'UL2: Unifying Language Learning Paradigms', Tay et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Can Language Models Learn from Explanations in Context?”, Lampinen et al 2022`{=html} {#lampinen-et-al-2022-section}

**[`Can language models learn from explanations in context?`{=html}](https://arxiv.org/abs/2204.02329#deepmind "'Can language models learn from explanations in context?', Lampinen et al 2022"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Socratic Models: Composing Zero-Shot Multimodal Reasoning With Language”, Zeng et al 2022`{=html} {#zeng-et-al-2022-2-section}

**[`Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language`{=html}](https://arxiv.org/abs/2204.00598#google "'Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language', Zeng et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“STaR: Bootstrapping Reasoning With Reasoning”, Zelikman et al 2022`{=html} {#zelikman-et-al-2022-section}

**[`STaR: Bootstrapping Reasoning With Reasoning`{=html}](https://arxiv.org/abs/2203.14465 "'STaR: Bootstrapping Reasoning With Reasoning', Zelikman et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“A Conversational Paradigm for Program Synthesis”, Nijkamp et al 2022`{=html} {#nijkamp-et-al-2022-2-section}

**[`A Conversational Paradigm for Program Synthesis`{=html}](https://arxiv.org/abs/2203.13474#salesforce "'A Conversational Paradigm for Program Synthesis', Nijkamp et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Self-Consistency Improves Chain-Of-Thought Reasoning in Language Models”, Wang et al 2022`{=html} {#wang-et-al-2022-20-section}

**[`Self-Consistency Improves Chain-of-Thought Reasoning in Language Models`{=html}](https://arxiv.org/abs/2203.11171#google "'Self-Consistency Improves Chain-of-Thought Reasoning in Language Models', Wang et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Learning-By-Narrating: Narrative Pre-Training for Zero-Shot Dialogue Comprehension”, Zhao et al 2022`{=html} {#zhao-et-al-2022-2-section}

**[`Learning-by-Narrating: Narrative Pre-Training for Zero-Shot Dialogue Comprehension`{=html}](https://arxiv.org/abs/2203.10249 "'Learning-by-Narrating: Narrative Pre-Training for Zero-Shot Dialogue Comprehension', Zhao et al 2022"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“PromptChainer: Chaining Large Language Model Prompts through Visual Programming”, Wu et al 2022`{=html} {#wu-et-al-2022-09-section}

**[`PromptChainer: Chaining Large Language Model Prompts through Visual Programming`{=html}](https://arxiv.org/abs/2203.06566#google "'PromptChainer: Chaining Large Language Model Prompts through Visual Programming', Wu et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models”, Wei et al 2022`{=html} {#wei-et-al-2022-4-section}

**[`Chain-of-Thought Prompting Elicits Reasoning in Large Language Models`{=html}](https://arxiv.org/abs/2201.11903#google "'Chain-of-Thought Prompting Elicits Reasoning in Large Language Models', Wei et al 2022"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Reasoning Like Program Executors”, Pi et al 2022`{=html} {#pi-et-al-2022-section}

**[`Reasoning Like Program Executors`{=html}](https://arxiv.org/abs/2201.11473#microsoft "'Reasoning Like Program Executors', Pi et al 2022"){.link-annotated
.id-not .include-annotation link-icon="MS"
link-icon-type="text,sans,italic"}**

## `“A Neural Network Solves and Generates Mathematics Problems by Program Synthesis: Calculus, Differential Equations, Linear Algebra, and More”, Drori et al 2021`{=html} {#drori-et-al-2021-section}

**[`A Neural Network Solves and Generates Mathematics Problems by Program Synthesis: Calculus, Differential Equations, Linear Algebra, and More`{=html}](https://arxiv.org/abs/2112.15594 "'A Neural Network Solves and Generates Mathematics Problems by Program Synthesis: Calculus, Differential Equations, Linear Algebra, and More', Drori et al 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“DREAM: Uncovering Mental Models behind Language Models”, Gu et al 2021`{=html} {#gu-et-al-2021-1-section}

**[`DREAM: Uncovering Mental Models behind Language Models`{=html}](https://arxiv.org/abs/2112.08656#allen "'DREAM: Uncovering Mental Models behind Language Models', Gu et al 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Reframing Human-AI Collaboration for Generating Free-Text Explanations”, Wiegreffe et al 2021`{=html} {#wiegreffe-et-al-2021-section}

**[`Reframing Human-AI Collaboration for Generating Free-Text Explanations`{=html}](https://arxiv.org/abs/2112.08674#allen "'Reframing Human-AI Collaboration for Generating Free-Text Explanations', Wiegreffe et al 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“NeuroLogic A<sup>✱</sup>esque Decoding: Constrained Text Generation With Lookahead Heuristics”, Lu et al 2021`{=html} {#lu-et-al-2021-1-section}

**[`NeuroLogic A<sup>✱</sup>esque Decoding: Constrained Text Generation with Lookahead Heuristics`{=html}](https://arxiv.org/abs/2112.08726#allen "'NeuroLogic A<sup>✱</sup>esque Decoding: Constrained Text Generation with Lookahead Heuristics', Lu et al 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“WebGPT: Improving the Factual Accuracy of Language Models through Web Browsing”, Hilton et al 2021`{=html} {#hilton-et-al-2021-1-section}

**[`WebGPT: Improving the factual accuracy of language models through web browsing`{=html}](https://openai.com/research/webgpt "'WebGPT: Improving the factual accuracy of language models through web browsing', Hilton et al 2021"){.link-annotated
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“NN Inner Monologue”, Gwern 2021`{=html} {#gwern-doc-ai-nn-transformer-gpt-inner-monologue-abstract-section}

**[`NN Inner Monologue`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/abstract "'NN Inner Monologue', Gwern 2021"){.link-annotated-partial
.id-not .include-annotation}**

## `“Few-Shot Self-Rationalization With Natural Language Prompts”, Marasović et al 2021`{=html} {#marasović-et-al-2021-section}

**[`Few-Shot Self-Rationalization with Natural Language Prompts`{=html}](https://arxiv.org/abs/2111.08284#allen "'Few-Shot Self-Rationalization with Natural Language Prompts', Marasović et al 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Training Verifiers to Solve Math Word Problems”, Cobbe et al 2021`{=html} {#cobbe-et-al-2021-section}

**[`Training Verifiers to Solve Math Word Problems`{=html}](https://arxiv.org/abs/2110.14168#openai "'Training Verifiers to Solve Math Word Problems', Cobbe et al 2021"){.link-annotated
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Unsupervised Neural Machine Translation With Generative Language Models Only”, Han et al 2021`{=html} {#han-et-al-2021-2-section}

**[`Unsupervised Neural Machine Translation with Generative Language Models Only`{=html}](https://arxiv.org/abs/2110.05448#openai "'Unsupervised Neural Machine Translation with Generative Language Models Only', Han et al 2021"){.link-annotated
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Show Your Work: Scratchpads for Intermediate Computation With Language Models”, Nye et al 2021`{=html} {#nye-et-al-2021-section}

**[`Show Your Work: Scratchpads for Intermediate Computation with Language Models`{=html}](https://arxiv.org/abs/2112.00114#google "'Show Your Work: Scratchpads for Intermediate Computation with Language Models', Nye et al 2021"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts”, Wu et al 2021`{=html} {#wu-et-al-2021-07-section}

**[`AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts`{=html}](https://arxiv.org/abs/2110.01691#google "'AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts', Wu et al 2021"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Teaching Autoregressive Language Models Complex Tasks By Demonstration”, Recchia 2021`{=html} {#recchia-2021-section}

**[`Teaching Autoregressive Language Models Complex Tasks By Demonstration`{=html}](https://arxiv.org/abs/2109.02102 "'Teaching Autoregressive Language Models Complex Tasks By Demonstration', Recchia 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Program Synthesis With Large Language Models”, Austin et al 2021`{=html} {#austin-et-al-2021-1-section}

**[`Program Synthesis with Large Language Models`{=html}](https://arxiv.org/abs/2108.07732#google "'Program Synthesis with Large Language Models', Austin et al 2021"){.link-annotated
.id-not .include-annotation link-icon="alphabet" link-icon-type="svg"
link-icon-color="#4285f4"}**

## `“Decision Transformer: Reinforcement Learning via Sequence Modeling”, Chen et al 2021`{=html} {#decisiontransformer-blog-section}

**[`Decision Transformer: Reinforcement Learning via Sequence Modeling`{=html}](https://sites.google.com/berkeley.edu/decision-transformer "'Decision Transformer: Reinforcement Learning via Sequence Modeling', Chen et al 2021"){.link-annotated
.id-not .include-annotation link-icon="BAIR"
link-icon-type="text,quad,mono"}**

## `“Explainable Multi-Hop Verbal Reasoning Through Internal Monologue”, Liang et al 2021`{=html} {#liang-et-al-2021-2-section}

**[`Explainable Multi-hop Verbal Reasoning Through Internal Monologue`{=html}](https://aclanthology.org/2021.naacl-main.97/ "'Explainable Multi-hop Verbal Reasoning Through Internal Monologue', Liang et al 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“A Simple Method to Keep GPT-3 Focused in a Conversation”, Mayne 2021`{=html} {#mayne-2021-section}

**[`A simple method to keep GPT-3 focused in a conversation`{=html}](https://andrewmayne.com/2021/05/18/a-simple-method-to-keep-gpt-3-focused-in-a-conversation/ "'A simple method to keep GPT-3 focused in a conversation', Mayne 2021"){.link-annotated
.id-not .include-annotation}**

## `“Measuring Mathematical Problem Solving With the MATH Dataset”, Hendrycks et al 2021`{=html} {#hendrycks-et-al-2021-4-section}

**[`Measuring Mathematical Problem Solving With the MATH Dataset`{=html}](https://arxiv.org/abs/2103.03874 "'Measuring Mathematical Problem Solving With the MATH Dataset', Hendrycks et al 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm”, Reynolds & McDonell 2021`{=html} {#reynolds-mcdonell-2021-section}

**[`Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm`{=html}](https://arxiv.org/abs/2102.07350 "'Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm', Reynolds & McDonell 2021"){.link-annotated
.id-not .include-annotation link-icon="𝛘" link-icon-type="text"
link-icon-color="#b31b1b"}**

## `“How We Accidentally Gave Our Bots Their Personalities”, Latitude 2021`{=html} {#latitude-2021-section}

**[`How We Accidentally Gave our Bots Their Personalities`{=html}](https://latitude.io/blog/how-we-accidentally-gave-our-bots-their-personalities/ "'How We Accidentally Gave our Bots Their Personalities', Latitude 2021"){.link-annotated
.id-not .include-annotation link-icon="AID"
link-icon-type="text,tri,sans"}**

## `“GPT-3: Imitation Learning That Imitates Learning”, Robertson 2020`{=html} {#robertson-2020-section}

**[`GPT-3: Imitation Learning that Imitates Learning`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2020-robertson.pdf "'GPT-3: Imitation Learning that Imitates Learning', Robertson 2020"){.link-annotated-partial
.id-not .include-annotation link-icon="pdf" link-icon-type="svg"
link-icon-color="#f40f02"}**

## `“Word in Context: Agent and Agent Clarification (69% Dev)”, Brockman 2020`{=html} {#brockman-2020-section}

**[`Word in Context: Agent and Agent Clarification (69% Dev)`{=html}](https://gptprompts.wikidot.com/linguistics:word-in-context#toc3 "'Word in Context: Agent and Agent Clarification (69% Dev)', Brockman 2020"){.link-annotated
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“I Found That Getting GPT-3 to Add Its Own "Internal Monologue" in Parentheses to Be a Helpful Strategy…”, blixt 2020`{=html} {#blixt-2020-section}

**[`I found that getting GPT-3 to add its own "internal monologue" in parentheses to be a helpful strategy…`{=html}](https://news.ycombinator.com/item?id=23990902 "'I found that getting GPT-3 to add its own "internal monologue" in parentheses to be a helpful strategy…', blixt 2020"){.link-annotated
.id-not .include-annotation link-icon="hacker-news" link-icon-type="svg"
link-icon-color="#f26522"}**

## `“You Can Probably Amplify GPT-3 Directly”, Robertson 2020`{=html} {#robertson-2020-section}

**[`You Can Probably Amplify GPT-3 Directly`{=html}](https://www.lesswrong.com/posts/Mzrs4MSi58ujBLbBG/you-can-probably-amplify-gpt3-directly "'You Can Probably Amplify GPT-3 Directly', Robertson 2020"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `kleptid @ "2020-07-17"`{=html} {#karyokleptid-2020-2-section}

**[`Seems to work`{=html}](https://x.com/kleptid/status/1284069270603866113 "'Seems to work', KaryoKleptid 2020"){.link-annotated
.id-not .include-annotation link-icon="twitter" link-icon-type="svg"
link-icon-color="#1da1f2"}**

## `kleptid @ "2020-07-17"`{=html} {#karyokleptid-2020-1-section}

**[`Teaching GPT-3 to do a brute force ‘for loop’ checking answers also seems to work`{=html}](https://x.com/kleptid/status/1284098635689611264 "'Teaching GPT-3 to do a brute force ‘for loop’ checking answers also seems to work', KaryoKleptid 2020"){.link-annotated
.id-not .include-annotation link-icon="twitter" link-icon-type="svg"
link-icon-color="#1da1f2"}**

## `“<span class=editorial>[More Early 4chan Inner-Monologue Examples]</span>”, Anonymous 2020`{=html} {#4chan-2020-inner-monologue-2-section}

**[`<span class=editorial>[More early 4chan inner-monologue examples]</span>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2020-07-16-4chan-vg-aidg-299570235-innermonologuexample-q299634976.html "'<span class=editorial>[More early 4chan inner-monologue examples]</span>', Anonymous 2020"){.link-annotated-partial
.id-not .include-annotation link-icon="internet-archive"
link-icon-type="svg"}**

## `“<span class=editorial>[4chan /vg/ Aidg Thread Screenshot of an Early Inner-Monologue for Arithmetic, Using a <em>Katawa Shoujo</em> Lilly Scenario]</span>”, Anonymous 2020`{=html} {#4chan-2020-inner-monologue-1-section}

**[`<span class=editorial>[4chan /vg/ aidg thread screenshot of an early inner-monologue for arithmetic, using a <em>Katawa Shoujo</em> Lilly scenario]</span>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2020-07-16-4chan-vg-aidg-anonymous-aidungeon2-katawashoujolilly-earlyinnermonologuepromptexample-1594872803390.png "'<span class=editorial>[4chan /vg/ aidg thread screenshot of an early inner-monologue for arithmetic, using a <em>Katawa Shoujo</em> Lilly scenario]</span>', Anonymous 2020"){.link-annotated
.id-not .include-annotation link-icon="image" link-icon-type="svg"}**

## `“[4chan /vg/ Board Discovers GPT-3 Inner-Monologues by Talking to the Wise Wolf Holo]”, Anonymous 2020`{=html} {#anonymous-2020-section}

**[`[4chan /vg/ board discovers GPT-3 inner-monologues by talking to the Wise Wolf Holo]`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2020-07-16-4chan-vg-aidungeon2thread-holoinnermonologue.html "'[4chan /vg/ board discovers GPT-3 inner-monologues by talking to the Wise Wolf Holo]', Anonymous 2020"){.link-annotated-partial
.id-not .include-annotation link-icon="internet-archive"
link-icon-type="svg"}**

## `“Inducing Self-Explanation: a Meta-Analysis”, Bisra et al 2018`{=html} {#bisra-et-al-2018-section}

**[`Inducing Self-Explanation: a Meta-Analysis`{=html}](/doc/psychology/spaced-repetition/2018-bisra.pdf "'Inducing Self-Explanation: a Meta-Analysis', Bisra et al 2018"){.link-annotated
.id-not .include-annotation link-icon="pdf" link-icon-type="svg"
link-icon-color="#f40f02"}**

## `“Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems”, Ling et al 2017`{=html} {#ling-et-al-2017-section}

**[`Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems`{=html}](https://arxiv.org/abs/1705.04146#deepmind "'Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems', Ling et al 2017"){.link-annotated
.id-not .include-annotation link-icon="deepmind" link-icon-type="svg"
link-icon-color="#4185f4"}**

## `“Why Do Humans Reason? Arguments for an Argumentative Theory”, Mercier & Sperber 2011`{=html} {#mercier-sperber-2011-section}

**[`Why do humans reason? Arguments for an argumentative theory`{=html}](https://hal.science/hal-00904097/document#pdf "'Why do humans reason? Arguments for an argumentative theory', Mercier & Sperber 2011"){.link-annotated-partial
.id-not .include-annotation}**

## `“Sebastian Riedel Homepage”, Riedel 2026`{=html} {#_rDMPMyzV-section}

**[`Sebastian Riedel homepage`{=html}](http://www.riedelcastro.org/ "'Sebastian Riedel homepage', Riedel 2026"){.link-annotated-partial
.id-not .include-annotation}**

## `“How to Dramatically Improve the Reasoning Ability of GPT-3”`{=html} {#_pTPg7vb5-section}

**[`How to dramatically improve the reasoning ability of GPT-3`{=html}](https://blog.andrewcantino.com/blog/2021/05/28/how-to-dramatically-improve-the-reasoning-ability-of-GPT-3/ "How to dramatically improve the reasoning ability of GPT-3"){.link-annotated-partial
.id-not .include-annotation}**

## `“A Preliminary Exploration into Factored Cognition With Language Models”`{=html} {#_13xdJAsk-section}

**[`A Preliminary Exploration into Factored Cognition with Language Models`{=html}](https://blog.eleuther.ai/factored-cognition/ "A Preliminary Exploration into Factored Cognition with Language Models"){.id-not
.include-annotation link-icon="eleutherai" link-icon-type="svg"}**

::: aux-links-transclude-file
::: collapse
**View External Link**:

[`https://blog.eleuther.ai/factored-cognition/`](https://blog.eleuther.ai/factored-cognition/){.id-not
.link-annotated-not .include-content .include-lazy
link-icon="eleutherai" link-icon-type="svg"}
:::
:::

## `“ChatGPT-4 O1-Pro: Poetry Reflection and Analysis”`{=html} {#_bW5NUard-section}

**[`ChatGPT-4 o1-pro: Poetry Reflection and Analysis`{=html}](https://chatgpt.com/share/67c101e3-2540-8006-b5d7-b742fa6f6936 "ChatGPT-4 o1-pro: Poetry Reflection and Analysis"){.link-annotated-partial
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“WiC_SelfContextStuffingImproved_Last10_stuft_examplesNV.ipynb”`{=html} {#_aZe9BdZM-section}

**[`WiC_SelfContextStuffingImproved_Last10_stuft_examplesNV.ipynb`{=html}](https://gist.github.com/brockmanmatt/deafb4dba7e4399327e44f2c8fd97b2b "WiC_SelfContextStuffingImproved_Last10_stuft_examplesNV.ipynb"){.id-not
.include-annotation link-icon="github" link-icon-type="svg"}**

## `“TinyZero”, Pan 2026`{=html} {#_y8xqbJui-section}

**[`TinyZero`{=html}](https://github.com/Jiayi-Pan/TinyZero "'TinyZero', Pan 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="github" link-icon-type="svg"}**

## `“Position Bias: A Benchmark for Testing Whether LLM Judges Keep the Same Preference When Two Lightly Edited Versions of the Same Story Are Shown in opposite Orders”, Mazir 2026`{=html} {#_-RnU-LLJ-section}

**[`Position bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders`{=html}](https://github.com/lechmazur/position_bias "'Position bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders', Mazir 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="github" link-icon-type="svg"}**

## `“Vincent-163/transformer-Arithmetic”`{=html} {#_XWN4ZoGh-section}

**[`vincent-163/transformer-arithmetic`{=html}](https://github.com/vincent-163/transformer-arithmetic "vincent-163/transformer-arithmetic"){.id-not
.include-annotation link-icon="github" link-icon-type="svg"}**

## `“Magic ToDo List Creator”`{=html} {#_7prboW4l-section}

**[`Magic ToDo List Creator`{=html}](https://goblin.tools/ "Magic ToDo List Creator"){.link-annotated-partial
.id-not .include-annotation}**

## `“Many Benchmarks Scores Would Appear Much Higher If You Let The AIs Use Adequate Labor”`{=html} {#_9_knqR98-section}

**[`Many Benchmarks Scores Would Appear Much Higher If You Let The AIs Use Adequate Labor`{=html}](https://joelbkr.substack.com/p/many-benchmarks-scores-would-appear "Many Benchmarks Scores Would Appear Much Higher If You Let The AIs Use Adequate Labor"){.link-annotated-partial
.id-not .include-annotation link-icon="substack" link-icon-type="svg"
link-icon-color="#ff6719"}**

## `“Short Story on AI: ‘Forward Pass’”, Karpathy 2026`{=html} {#_E2b7enKk-section}

**[`Short Story on AI: ‘Forward Pass’`{=html}](https://karpathy.github.io/2021/03/27/forward-pass/ "'Short Story on AI: ‘Forward Pass’', Karpathy 2026"){.id-not
.include-annotation}**

::: aux-links-transclude-file
::: collapse
**View External Link**:

[`https://karpathy.github.io/2021/03/27/forward-pass/`](https://karpathy.github.io/2021/03/27/forward-pass/){.id-not
.link-annotated-not .include-content .include-lazy}
:::
:::

## `“AI Dungeon Players Can Now Translate Their Stories into Emojis by Just Clicking a Button.”`{=html} {#_ZtmO6kUZ-section}

**[`AI Dungeon players can now translate their stories into emojis by just clicking a button.`{=html}](https://latitude.io/blog/introducing-ai-dungeon-translate/ "AI Dungeon players can now translate their stories into emojis by just clicking a button."){.link-annotated-partial
.id-not .include-annotation link-icon="AID"
link-icon-type="text,tri,sans"}**

## `“Sky-T1: Train Your Own <code>o1-Preview</code> Model With $450”`{=html} {#_Q1J5CvIG-section}

**[`Sky-T1: Train your own <code>o1-preview</code> model with $450`{=html}](https://novasky-ai.github.io/posts/sky-t1/ "Sky-T1: Train your own <code>o1-preview</code> model with $450"){.link-annotated-partial
.id-not .include-annotation}**

## `“Solving Math Word Problems: We’ve Trained a System That Solves Grade School Math Problems With Nearly Twice the Accuracy of a Fine-Tuned GPT-3 Model. It Solves about 90% As Many Problems As Real Kids: a Small Sample of 9-12 Year Olds Scored 60% on a Test from Our Dataset, While Our System Scored 55% on Those Same Problems. This Is Important Because Today’s AI Is Still Quite Weak at Commonsense Multistep Reasoning, Which Is Easy Even for Grade School Kids. We Achieved These Results by Training Our Model to Recognize Its Mistakes, so That It Can Try Repeatedly Until It Finds a Solution That Works”`{=html} {#_c7ujB62B-section}

**[`Solving Math Word Problems: We’ve trained a system that solves grade school math problems with nearly twice the accuracy of a fine-tuned GPT-3 model. It solves about 90% as many problems as real kids: a small sample of 9-12 year olds scored 60% on a test from our dataset, while our system scored 55% on those same problems. This is important because today’s AI is still quite weak at commonsense multistep reasoning, which is easy even for grade school kids. We achieved these results by training our model to recognize its mistakes, so that it can try repeatedly until it finds a solution that works`{=html}](https://openai.com/research/solving-math-word-problems "Solving Math Word Problems: We’ve trained a system that solves grade school math problems with nearly twice the accuracy of a fine-tuned GPT-3 model. It solves about 90% as many problems as real kids: a small sample of 9-12 year olds scored 60% on a test from our dataset, while our system scored 55% on those same problems. This is important because today’s AI is still quite weak at commonsense multistep reasoning, which is easy even for grade school kids. We achieved these results by training our model to recognize its mistakes, so that it can try repeatedly until it finds a solution that works"){.id-not
.include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Prompting Diverse Ideas: Increasing AI Idea Variance”`{=html} {#_2N5zAcx0-section}

**[`Prompting Diverse Ideas: Increasing AI Idea Variance`{=html}](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4708466 "Prompting Diverse Ideas: Increasing AI Idea Variance"){.link-annotated-partial
.id-not .include-annotation link-icon="SSRN" link-icon-type="text,quad"
link-icon-color="#007398"}**

## `“<code>o3-Mini</code>”, OpenAI 2026`{=html} {#_RTv5vPLd-section}

**[`<code>o3-mini</code>`{=html}](https://platform.openai.com/docs/models#o3-mini "'<code>o3-mini</code>', OpenAI 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="openai" link-icon-type="svg"}**

## `“Teaching a Neural Network to Use a Calculator”`{=html} {#_IvpL9_Bq-section}

**[`Teaching a neural network to use a calculator`{=html}](https://reiinakano.com/2019/11/12/solving-probability.html "Teaching a neural network to use a calculator"){.link-annotated-partial
.id-not .include-annotation}**

## `“Can Tiny Language Models Reason? [Inner-Monologue & DPO RLHF on a 0.13b-Parameter LLM: <code>trlm</code>]”`{=html} {#_IkDXs1l6-section}

**[`Can Tiny Language Models Reason? [inner-monologue & DPO RLHF on a 0.13b-parameter LLM: <code>trlm</code>]`{=html}](https://shekswess.github.io/tiny-reasoning-language-model.html "Can Tiny Language Models Reason? [inner-monologue & DPO RLHF on a 0.13b-parameter LLM: <code>trlm</code>]"){.id-not
.include-annotation}**

::: aux-links-transclude-file
::: collapse
**View HTML**:

[`https://shekswess.github.io/tiny-reasoning-language-model.html`](https://shekswess.github.io/tiny-reasoning-language-model.html){.id-not
.link-annotated-not .include-content .include-lazy}
:::
:::

## `“On-Policy Distillation”`{=html} {#_Jdz9b3YH-section}

**[`On-Policy Distillation`{=html}](https://thinkingmachines.ai/blog/on-policy-distillation/#distillation-for-personalization "On-Policy Distillation"){.link-annotated-partial
.id-not .include-annotation}**

## `“GPT-4 O1 Isn’t a Chat Model (And That’s the Point)”`{=html} {#_qUMxFNhN-section}

**[`GPT-4 o1 isn’t a chat model (and that’s the point)`{=html}](https://www.latent.space/p/o1-skill-issue "GPT-4 o1 isn’t a chat model (and that’s the point)"){.link-annotated-partial
.id-not .include-annotation}**

## `“Connecting the Dots: LLMs Can Infer &amp; Verbalize Latent Structure from Training Data”`{=html} {#_yCsWDcTM-section}

**[`Connecting the Dots: LLMs can Infer &amp; Verbalize Latent Structure from Training Data`{=html}](https://www.lesswrong.com/posts/5SKRHQEFr8wYQHYkx/connecting-the-dots-llms-can-infer-and-verbalize-latent "Connecting the Dots: LLMs can Infer &amp; Verbalize Latent Structure from Training Data"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Beware General Claims about ‘Generalizable Reasoning Capabilities’ (Of Modern AI Systems)”`{=html} {#_t8w9kEcz-section}

**[`Beware General Claims about ‘Generalizable Reasoning Capabilities’ (of Modern AI Systems)`{=html}](https://www.lesswrong.com/posts/5uw26uDdFbFQgKzih/beware-general-claims-about-generalizable-reasoning "Beware General Claims about ‘Generalizable Reasoning Capabilities’ (of Modern AI Systems)"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Preventing Language Models from Hiding Their Reasoning”`{=html} {#_SRnWhNo8-section}

**[`Preventing Language Models from hiding their reasoning`{=html}](https://www.lesswrong.com/posts/9Fdd9N7Escg3tcymb/preventing-language-models-from-hiding-their-reasoning "Preventing Language Models from hiding their reasoning"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Claude Opus 4.6 Reasoning Doesn’t Verbalize Alignment Faking, but Behavior Persists”`{=html} {#_FIigrOSU-section}

**[`Claude Opus 4.6 Reasoning Doesn’t Verbalize Alignment Faking, but Behavior Persists`{=html}](https://www.lesswrong.com/posts/9wDHByRhmtDaoYAx8/opus-4-6-reasoning-doesn-t-verbalize-alignment-faking-but "Claude Opus 4.6 Reasoning Doesn’t Verbalize Alignment Faking, but Behavior Persists"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“A High Level Closed-Door Session Discussing DeepSeek: Vision Trumps Technology”`{=html} {#_grCB6vBe-section}

**[`A High Level Closed-Door Session Discussing DeepSeek: Vision Trumps Technology`{=html}](https://www.lesswrong.com/posts/JTKaR5q59BgDp6rH8/a-high-level-closed-door-session-discussing-deepseek-vision "A High Level Closed-Door Session Discussing DeepSeek: Vision Trumps Technology"){.link-annotated-partial
.id-not .include-annotation link-icon="deepseek" link-icon-type="svg"
link-icon-color="#4d6bfe"}**

## `“Do Models Continue Misaligned Actions?”`{=html} {#_KrYpgOX2-section}

**[`Do Models Continue Misaligned Actions?`{=html}](https://www.lesswrong.com/posts/SawczP2pdCXMrkg2A/do-models-continue-misaligned-actions-2 "Do Models Continue Misaligned Actions?"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“How Well Do Models Follow Their Constitutions?”`{=html} {#_YOy9zUnM-section}

**[`How well do models follow their constitutions?`{=html}](https://www.lesswrong.com/posts/Tk4SF8qFdMrzGJGGw/how-well-do-models-follow-their-constitutions "How well do models follow their constitutions?"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Opus’s Schelling Steganography Has Amplifiable Secrecy Against Weaker Eavesdroppers”`{=html} {#_as6ahmPk-section}

**[`Opus’s Schelling Steganography Has Amplifiable Secrecy Against Weaker Eavesdroppers`{=html}](https://www.lesswrong.com/posts/e5sdgYxYFKqdM6vF4/opus-s-schelling-steganography-has-amplifiable-secrecy "Opus’s Schelling Steganography Has Amplifiable Secrecy Against Weaker Eavesdroppers"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Did Claude 3 Opus Align Itself via Gradient Hacking?”`{=html} {#_k4iqFsKE-section}

**[`Did Claude 3 Opus align itself via gradient hacking?`{=html}](https://www.lesswrong.com/posts/ioZxrP7BhS5ArK59w/did-claude-3-opus-align-itself-via-gradient-hacking "Did Claude 3 Opus align itself via gradient hacking?"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“What Secret Goals Does Claude Think It Has?”`{=html} {#_1pMJYe7u-section}

**[`What secret goals does Claude think it has?`{=html}](https://www.lesswrong.com/posts/mYM9EAAhpbYDDmA3e/what-secret-goals-does-claude-think-it-has "What secret goals does Claude think it has?"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Monitor Jailbreaking: Evading Chain-Of-Thought Monitoring Without Encoded Reasoning”`{=html} {#_TgrN39Fq-section}

**[`Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning`{=html}](https://www.lesswrong.com/posts/szyZi5d4febZZSiq3/monitor-jailbreaking-evading-chain-of-thought-monitoring "Monitor Jailbreaking: Evading Chain-of-Thought Monitoring Without Encoded Reasoning"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Steganography in Chain-Of-Thought Reasoning”`{=html} {#_JucNT34_-section}

**[`Steganography in Chain-of-Thought Reasoning`{=html}](https://www.lesswrong.com/posts/yDcMDJeSck7SuBs24/steganography-in-chain-of-thought-reasoning "Steganography in Chain-of-Thought Reasoning"){.link-annotated-partial
.id-not .include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

## `“Visible Thoughts Project and Bounty Announcement”`{=html} {#_qQHbVo4d-section}

**[`Visible Thoughts Project and Bounty Announcement`{=html}](https://www.lesswrong.com/posts/zRn6cLtxyNodudzhw/visible-thoughts-project-and-bounty-announcement "Visible Thoughts Project and Bounty Announcement"){.id-not
.include-annotation link-icon="LW" link-icon-type="text"
link-icon-color="#7faf83"}**

::: aux-links-transclude-file
::: collapse
**View External Link**:

[`https://www.lesswrong.com/posts/zRn6cLtxyNodudzhw/visible-thoughts-project-and-bounty-announcement`](https://www.lesswrong.com/posts/zRn6cLtxyNodudzhw/visible-thoughts-project-and-bounty-announcement){.id-not
.link-annotated-not .include-content .include-lazy link-icon="LW"
link-icon-type="text" link-icon-color="#7faf83"}
:::
:::

## `Malcolm_Ocean`{=html} {#_tqLUSsgH-section}

**[`Inspired by an AI Dungeon example where math is discussed in simple language, I seem to be having decent results here. I had to… not just say what parity IS but HOW to calculate it (‘count the number of 1s’) and then it sort of walks itself through decently. Tho kinda confused`{=html}](https://x.com/Malcolm_Ocean/status/1285099206781341696 "'Inspired by an AI Dungeon example where math is discussed in simple language, I seem to be having decent results here. I had to… not just say what parity IS but HOW to calculate it (‘count the number of 1s’) and then it sort of walks itself through decently. Tho kinda confused', Ocean 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="twitter" link-icon-type="svg"
link-icon-color="#1da1f2"}**

## `bucketofkets`{=html} {#_SUlHZLvn-section}

**[`I think ‘GPT-3 can’t do parity checking’ isn’t quite right. It can clearly pattern match the algorithm, almost perfectly. It’s just a little mistake prone. Here, I invented a syntax for having it evaluate parity on each pair of digits. It…almost gets it right.`{=html}](https://x.com/bucketofkets/status/1285100951271952384 "'I think ‘GPT-3 can’t do parity checking’ isn’t quite right. It can clearly pattern match the algorithm, almost perfectly. It’s just a little mistake prone. Here, I invented a syntax for having it evaluate parity on each pair of digits. It…almost gets it right.', bucketofkets 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="twitter" link-icon-type="svg"
link-icon-color="#1da1f2"}**

## `sama`{=html} {#_BfXo22Z5-section}

**[`[o3-full & o4-mini to launch earlier, GPT-5 delayed for capability improvement, integration polishing, &amp; hardware availability]`{=html}](https://x.com/sama/status/1908167621624856998 "'[o3-full & o4-mini to launch earlier, GPT-5 delayed for capability improvement, integration polishing, &amp; hardware availability]', Altman 2026"){.link-annotated-partial
.id-not .include-annotation link-icon="twitter" link-icon-type="svg"
link-icon-color="#1da1f2"}**

## `teortaxesTex`{=html} {#_gKc8xZvg-section}

**[`[DeepSeek-r1 solving Russian pun]`{=html}](https://x.com/teortaxesTex/status/1879874030372802930 "'[DeepSeek-r1 solving Russian pun]', teortaxesTex 2026"){.id-not
.include-annotation link-icon="twitter" link-icon-type="svg"
link-icon-color="#1da1f2"}**

::: aux-links-transclude-file
[`https://x.com/teortaxesTex/status/1879874030372802930`](https://x.com/teortaxesTex/status/1879874030372802930 "'[DeepSeek-r1 solving Russian pun]', teortaxesTex 2026"){.id-not
.include-content}
:::

## Sort By Magic

::: {demo-type="sort-by-magic-preface"}
Annotations sorted by machine learning into [inferred \'tags\'](/design#future-tag-features){#gwern-design--future-tag-features
.link-annotated-partial}. This provides an alternative way to browse: instead of by *date* order, one can browse in *topic* order. The \'sorted\' list has been automatically clustered into multiple sections & auto-labeled for easier browsing.

Beginning with the newest annotation, it uses the embedding of each annotation to attempt to create a list of nearest-neighbor annotations, creating a progression of topics. For more details, see the link.
:::

### `reasoning-rate android-dataset control-system low-utilization android-devices` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#rawles-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#altman-2025-section){.include
.include-even-when-collapsed}

### `test-driven-generation` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#han-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#chen-et-al-2022-codet-section){.include
.include-even-when-collapsed}

### `gpt-summarization-diffusion` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#adams-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#engels-et-al-2026-section){.include
.include-even-when-collapsed}

### `agent-scaling` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#li-et-al-2024-10-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#gan-isola-2026-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zheng-et-al-2025-section){.include
.include-even-when-collapsed}

### `unsupervised-translation` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#ma-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#han-et-al-2021-2-section){.include
.include-even-when-collapsed}

### `tabular-gpt` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#ye-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zha-et-al-2023-section){.include
.include-even-when-collapsed}

### `lateral-thinking-reranking` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#todd-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lee-et-al-2024-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#weller-et-al-2025-section){.include
.include-even-when-collapsed}

### `training-efficiency` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#chen-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#agarwal-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hsieh-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#snell-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#nye-et-al-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#karyokleptid-2020-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#karyokleptid-2020-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#4chan-2020-inner-monologue-1-section){.include
.include-even-when-collapsed}

### `reasoning-medical jailbreaks steganography finance-prediction generalist-models` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#liévin-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#nori-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#callanan-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#pham-cunningham-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zolkowski-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#russinovich-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#mehrotra-et-al-2023-section){.include
.include-even-when-collapsed}

### `chain-of-thought` {.collapse title="Machine-generated tag name for the following cluster of links."}

[\[see previous entry\]](#_-RnU-LLJ-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#karvonen-marks-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#roger-greenblatt-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#guo-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#turpin-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lanham-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lyu-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#chua-evans-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#creswell-shanahan-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#chen-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#feng-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#li-et-al-2024-09-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#merrill-sabharwal-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#stechly-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-zhou-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#sprague-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#liu-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#jin-et-al-2024-4-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#cheng-durme-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hao-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#su-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#pfau-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#treutlein-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#deng-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#deng-et-al-2024-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#phan-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wei-et-al-2022-4-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2022-20-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#press-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#tay-et-al-2022-ul2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#tay-et-al-2022-upalm-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#paliotta-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#geiping-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#muennighoff-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#yuan-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ye-et-al-2024-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#qian-et-al-2022-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#bai-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ramesh-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#abbe-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hahn-rofin-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lee-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#singh-strouse-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#mccoy-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#mccoy-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#levy-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#anil-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhou-et-al-2022-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#kojima-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhang-et-al-2023-19-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#shi-et-al-2022-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#gao-et-al-2022-4-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#schlag-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#liu-et-al-2023-16-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#valmeekam-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhou-et-al-2023-04-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#yao-et-al-2022-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#mezghani-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hu-clune-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wu-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#yao-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#besta-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ding-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wu-et-al-2024-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#huang-et-al-2023-7-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wong-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#gu-et-al-2021-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#moghaddam-honey-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2023-12-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#deng-et-al-2023-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#xu-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#leviathan-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#reynolds-mcdonell-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wu-et-al-2022-09-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wu-et-al-2021-07-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#dohan-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#chen-et-al-2023-11-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#pilault-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lampinen-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wiegreffe-et-al-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#reppert-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#radhakrishnan-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#xie-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lee-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hosseini-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zelikman-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zelikman-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#goyal-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wu-et-al-2024-4-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wu-et-al-2024-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wen-et-al-2024-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#delétang-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2024-07-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#kumar-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#huang-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#pan-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lu-et-al-2024-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#shinn-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hong-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhao-et-al-2022-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#dai-et-al-2022-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#nijkamp-et-al-2022-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#austin-et-al-2021-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#haluptzok-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#huang-et-al-2022-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#singh-et-al-2023-3-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#dang-ngo-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#guo-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#cheng-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#yue-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#karan-du-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ruis-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ward-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#pi-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#liang-et-al-2021-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#jung-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ma-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#arnav-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#korbak-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#bowman-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#needham-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#kadavath-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#nakkiran-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhang-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#xu-zhang-2026-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#du-et-al-2023-3-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wei-et-al-2024-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#kim-et-al-2024-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#chakrabarty-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ankner-et-al-2024-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#yuan-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#luo-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lightman-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#uesato-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hendrycks-et-al-2021-4-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#yuan-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhou-et-al-2023-06-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhou-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#cobbe-et-al-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lu-et-al-2022-3-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#drori-et-al-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lewkowycz-et-al-2022-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#recchia-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#honovich-et-al-2022-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#kim-et-al-2023-6-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wu-et-al-2023-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ifargan-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#narayanan-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#phan-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#gerrits-2026-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#aiken-2026-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#el-kishky-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#li-et-al-2023-05-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ling-et-al-2017-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#morris-et-al-2026-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#mukherjee-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ji-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#decisiontransformer-blog-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#huang-et-al-2022-5-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zeng-et-al-2022-2-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#chen-et-al-2023-03-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#saha-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#obrien-lewis-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#lu-et-al-2021-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#deepseek-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#huang-yang-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#huang-et-al-2024-3-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhong-et-al-2024-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#martínez-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#choi-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#nay-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#nay-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#brockman-2020-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#latitude-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#blixt-2020-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#mayne-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#hilton-et-al-2021-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#manakul-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#zhang-et-al-2023-14-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#scheurer-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#meinke-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#payne-alloui-cros-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#li-shirado-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#liang-et-al-2025-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#parcalabescu-frank-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#marasović-et-al-2021-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#bisra-et-al-2018-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#manikandan-et-al-2023-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#ruan-et-al-2024-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#suzgun-et-al-2022-1-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#wang-et-al-2024-06-section){.include
.include-even-when-collapsed}

[\[see previous entry\]](#openai-2025-section){.include
.include-even-when-collapsed}

## Wikipedia (4) {#titled-links-wikipedia .collapse}

1.  **[`Language of thought hypothesis`{=html}](https://en.wikipedia.org/wiki/Language_of_thought_hypothesis "Language of thought hypothesis"){.link-annotated-partial
    .id-not .include-annotation link-icon="wikipedia"
    link-icon-type="svg"}**

2.  **[`Retrospective think aloud`{=html}](https://en.wikipedia.org/wiki/Retrospective_think_aloud "Retrospective think aloud"){.id-not
    .include-annotation link-icon="wikipedia" link-icon-type="svg"}**

    ::: aux-links-transclude-file
    [`https://en.wikipedia.org/wiki/Retrospective_think_aloud`](https://en.wikipedia.org/wiki/Retrospective_think_aloud "Retrospective think aloud"){.id-not
    .include-content include-template="$annotationFileIncludeTemplate"}
    :::

3.  **[`Rubber duck debugging`{=html}](https://en.wikipedia.org/wiki/Rubber_duck_debugging "Rubber duck debugging"){.id-not
    .include-annotation link-icon="wikipedia" link-icon-type="svg"}**

    ::: aux-links-transclude-file
    [`https://en.wikipedia.org/wiki/Rubber_duck_debugging`](https://en.wikipedia.org/wiki/Rubber_duck_debugging "Rubber duck debugging"){.id-not
    .include-content include-template="$annotationFileIncludeTemplate"}
    :::

4.  **[`Think aloud protocol`{=html}](https://en.wikipedia.org/wiki/Think_aloud_protocol "Think aloud protocol"){.link-annotated-partial
    .id-not .include-annotation link-icon="wikipedia"
    link-icon-type="svg"}**

# Miscellaneous

-   **[`<code>https://x.com/Thomas_Woodside/status/1936284255426167073</code>`{=html}](https://x.com/Thomas_Woodside/status/1936284255426167073 "Thomas_Woodside 2025"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-chen-table1-gpt35promptsusedtorepeatedlyrefinenaturallanguagetranslationsinnermonologuestyle.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-chen-table1-gpt35promptsusedtorepeatedlyrefinenaturallanguagetranslationsinnermonologuestyle.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure1-numberformattingforgpt2arithmetic.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure1-numberformattingforgpt2arithmetic.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure2-thefourinputformattingoptionsforgptinnermonologue.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure2-thefourinputformattingoptionsforgptinnermonologue.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure3-performanceofgpton3digitarithmeticdependsondatadistribution.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure3-performanceofgpton3digitarithmeticdependsondatadistribution.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure9-arithmeticcanbelearnedevenwithnoiseintheinnermonologuetranscripts.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-lee-figure9-arithmeticcanbelearnedevenwithnoiseintheinnermonologuetranscripts.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-moghaddam-figure1-examplesofzerovstwoshottheoryofmindprompting.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-moghaddam-figure1-examplesofzerovstwoshottheoryofmindprompting.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-moghaddam-figure3-gpt3andgpt4performanceontheoryofmindwithinnermonologues.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-moghaddam-figure3-gpt3andgpt4performanceontheoryofmindwithinnermonologues.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-pilaut-figure2-exampleambiguitiesintranslatingfrenchtoenglish.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-pilaut-figure2-exampleambiguitiesintranslatingfrenchtoenglish.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2023-pilaut-figure3-interceptinnermonologuequestionaskingonlyemergesatscalefrompalm62bto540b.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2023-pilaut-figure3-interceptinnermonologuequestionaskingonlyemergesatscalefrompalm62bto540b.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/reinforcement-learning/imitation-learning/2023-lee-figure6-sampleefficiencyofvariousinnermonologueformatsshowingmoredetailedisbetterforimitationlearning.png</code>`{=html}](/doc/reinforcement-learning/imitation-learning/2023-lee-figure6-sampleefficiencyofvariousinnermonologueformatsshowingmoredetailedisbetterforimitationlearning.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-10-24-raldi-gpt3doesanastonishinglygoodjobcreatingbothsidesofaninteractivefictiontranscript.html</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-10-24-raldi-gpt3doesanastonishinglygoodjobcreatingbothsidesofaninteractivefictiontranscript.html){.link-annotated-partial
    .id-not .include-annotation link-icon="internet-archive"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-05-28-gpt3user-thinkingisallyouneed.html</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-05-28-gpt3user-thinkingisallyouneed.html){.link-annotated-partial
    .id-not .include-annotation link-icon="internet-archive"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-dai-figure4-qreccretrevialperformancelogscalinginwikidialogdatasetsize.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-dai-figure4-qreccretrevialperformancelogscalinginwikidialogdatasetsize.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure2-3kindsofnaturallanguagefeedbackforcontrollingsaycaninnermonologue.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure2-3kindsofnaturallanguagefeedbackforcontrollingsaycaninnermonologue.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure3-testinginnermonologuein3roboticdomains.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure3-testinginnermonologuein3roboticdomains.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5a-emergentcapabilities-continuedadaptationtonewinstructions.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5a-emergentcapabilities-continuedadaptationtonewinstructions.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5b-emergentcapabilities-selfproposingnewgoalsunderinfeasibilityofoldgoals.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5b-emergentcapabilities-selfproposingnewgoalsunderinfeasibilityofoldgoals.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5c-emergentcapabilities-multilingualinteractioninchinese.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5c-emergentcapabilities-multilingualinteractioninchinese.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5d-emergentcapabilities-interactivesceneunderstandinglikeshrdlu.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-huang-figure5d-emergentcapabilities-interactivesceneunderstandinglikeshrdlu.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-lampinen-figure2-gopherperformanceimprovementsfromexplanationofproblems.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-lampinen-figure2-gopherperformanceimprovementsfromexplanationofproblems.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-lampinen-figure4-largermodelsbenefitmorefromexplanationofproblems.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-lampinen-figure4-largermodelsbenefitmorefromexplanationofproblems.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-figure3-gpt3selfaskinnermonologuedemonstration.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-figure3-gpt3selfaskinnermonologuedemonstration.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-figure4-selfaskinnermonologueperformsequallywellon1hopand2hopquestionanswering.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-figure4-selfaskinnermonologueperformsequallywellon1hopand2hopquestionanswering.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-figure5-selfaskplusgooglesearchengine-innermonologueforsearchingtheinternettoanswermultihopquestions.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-figure5-selfaskplusgooglesearchengine-innermonologueforsearchingtheinternettoanswermultihopquestions.png){.link-annotated-partial
    .id-not .include-annotation link-icon="alphabet"
    link-icon-type="svg" link-icon-color="#4285f4"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-table1-selfaskplusgooglesearchengine-innermonologueforsearchingtheinternettoanswermultihopquestions-benchmarkperformance.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-press-table1-selfaskplusgooglesearchengine-innermonologueforsearchingtheinternettoanswermultihopquestions-benchmarkperformance.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="alphabet"
    link-icon-type="svg" link-icon-color="#4285f4"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-shi-figure4-multilingualinnermonologuescalingbyparametercountingpt3andpalm.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-shi-figure4-multilingualinnermonologuescalingbyparametercountingpt3andpalm.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-shi-figure5-multiglinalfewshotscalinginpalm540bbynumberofexamples.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-shi-figure5-multiglinalfewshotscalinginpalm540bbynumberofexamples.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-tay-ul2-innermonologueresults.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-tay-ul2-innermonologueresults.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-wang-figure2-selfconsistencycompletiongreatlyimprovesanswercorrectness.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-wang-figure2-selfconsistencycompletiongreatlyimprovesanswercorrectness.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure2-lamdamathwordproblemscalinginmodelparametersize.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure2-lamdamathwordproblemscalinginmodelparametersize.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure3-lamdamathwordproblemscalingwithmodelparametersizewhenusinginnermonologueprompts.jpg</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure3-lamdamathwordproblemscalingwithmodelparametersizewhenusinginnermonologueprompts.jpg){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure5-lamdamatsymbolicreasoningproblemscalingwithmodelparametersizewhenusinginnermonologueprompts.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure5-lamdamatsymbolicreasoningproblemscalingwithmodelparametersizewhenusinginnermonologueprompts.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure6-lamdacommonsensereasoningproblemscalingwithmodelparametersizewhenusinginnermonologueprompts.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure6-lamdacommonsensereasoningproblemscalingwithmodelparametersizewhenusinginnermonologueprompts.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure8-lamdavsgpt3.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-wei-figure8-lamdavsgpt3.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>/doc/ai/nn/transformer/gpt/inner-monologue/2022-zeng-figure2-socraticmodelsworkflowoverview.png</code>`{=html}](/doc/ai/nn/transformer/gpt/inner-monologue/2022-zeng-figure2-socraticmodelsworkflowoverview.png){.link-annotated-partial
    .id-not .include-annotation link-icon="image"
    link-icon-type="svg"}**

-   **[`<code>https://applied-llms.org/</code>`{=html}](https://applied-llms.org/){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://blog.valentin.sh/chatgpt5/</code>`{=html}](https://blog.valentin.sh/chatgpt5/){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://builtin.com/job/customer-success/expert-ai-teacher-contract/1267315</code>`{=html}](https://builtin.com/job/customer-success/expert-ai-teacher-contract/1267315){.id-not
    .include-annotation}**

-   **[`<code>https://curvy-check-498.notion.site/Process-Reinforcement-through-Implicit-Rewards-15f4fcb9c42180f1b498cc9b2eaf896f</code>`{=html}](https://curvy-check-498.notion.site/Process-Reinforcement-through-Implicit-Rewards-15f4fcb9c42180f1b498cc9b2eaf896f){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://generative.ink/posts/methods-of-prompt-programming/#serializing-reasoning</code>`{=html}](https://generative.ink/posts/methods-of-prompt-programming/#serializing-reasoning){.id-not
    .include-annotation}**

    ::: aux-links-transclude-file
    ::: collapse
    **View HTML**:

    [`/doc/www/generative.ink/2fa0cae05f887923f2a169fecaa094fa3075f6ba.html#serializing-reasoning`](/doc/www/generative.ink/2fa0cae05f887923f2a169fecaa094fa3075f6ba.html#serializing-reasoning){.id-not
    .link-annotated-not .include-content .include-lazy
    link-icon="internet-archive" link-icon-type="svg"}
    :::
    :::

-   **[`<code>https://github.com/OpenBioLink/ThoughtSource</code>`{=html}](https://github.com/OpenBioLink/ThoughtSource){.id-not
    .include-annotation link-icon="github" link-icon-type="svg"}**

-   **[`<code>https://github.com/desik1998/MathWithLLMs</code>`{=html}](https://github.com/desik1998/MathWithLLMs){.link-annotated-partial
    .id-not .include-annotation link-icon="github"
    link-icon-type="svg"}**

-   **[`<code>https://github.com/ggerganov/llama.cpp/pull/1773</code>`{=html}](https://github.com/ggerganov/llama.cpp/pull/1773){.link-annotated-partial
    .id-not .include-annotation link-icon="github"
    link-icon-type="svg"}**

-   **[`<code>https://github.com/openai/openai-cookbook/blob/main/techniques_to_improve_reliability.md</code>`{=html}](https://github.com/openai/openai-cookbook/blob/main/techniques_to_improve_reliability.md){.id-not
    .include-annotation link-icon="github" link-icon-type="svg"}**

-   **[`<code>https://jxnl.github.io/instructor/blog/2023/11/05/chain-of-density/</code>`{=html}](https://jxnl.github.io/instructor/blog/2023/11/05/chain-of-density/){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://lingo.csail.mit.edu/blog/arithmetic_gpt3/</code>`{=html}](https://lingo.csail.mit.edu/blog/arithmetic_gpt3/){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://model-checking.github.io/kani-verifier-blog/2023/05/01/writing-code-with-chatgpt-improve-it-with-kani.html</code>`{=html}](https://model-checking.github.io/kani-verifier-blog/2023/05/01/writing-code-with-chatgpt-improve-it-with-kani.html){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://niplav.site/decompose.html#Small_Experiment</code>`{=html}](https://niplav.site/decompose.html#Small_Experiment){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://openai.com/index/introducing-openai-o1-preview/</code>`{=html}](https://openai.com/index/introducing-openai-o1-preview/){.link-annotated-partial
    .id-not .include-annotation link-icon="openai"
    link-icon-type="svg"}**

-   **[`<code>https://platform.openai.com/docs/guides/reasoning/how-reasoning-works</code>`{=html}](https://platform.openai.com/docs/guides/reasoning/how-reasoning-works){.id-not
    .include-annotation link-icon="openai" link-icon-type="svg"}**

-   **[`<code>https://reasoning-tokens.ghost.io/reasoning-tokens/</code>`{=html}](https://reasoning-tokens.ghost.io/reasoning-tokens/){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://research.google/blog/google-research-2022-beyond-language-vision-and-generative-models/</code>`{=html}](https://research.google/blog/google-research-2022-beyond-language-vision-and-generative-models/){.link-annotated-partial
    .id-not .include-annotation link-icon="alphabet"
    link-icon-type="svg" link-icon-color="#4285f4"}**

-   **[`<code>https://research.google/blog/minerva-solving-quantitative-reasoning-problems-with-language-models/</code>`{=html}](https://research.google/blog/minerva-solving-quantitative-reasoning-problems-with-language-models/){.id-not
    .include-annotation link-icon="alphabet" link-icon-type="svg"
    link-icon-color="#4285f4"}**

-   **[`<code>https://statmodeling.stat.columbia.edu/2023/08/30/chatgpt-4-can-do-3-digit-multiplication/</code>`{=html}](https://statmodeling.stat.columbia.edu/2023/08/30/chatgpt-4-can-do-3-digit-multiplication/){.id-not
    .include-annotation link-icon="▅▇▃" link-icon-type="text,tri"}**

-   **[`<code>https://towardsdatascience.com/1-1-3-wait-no-1-1-2-how-to-have-gpt-sanity-check-itself-136e846987bf</code>`{=html}](https://towardsdatascience.com/1-1-3-wait-no-1-1-2-how-to-have-gpt-sanity-check-itself-136e846987bf){.link-annotated-partial
    .id-not .include-annotation link-icon="𝐌" link-icon-type="text"}**

-   **[`<code>https://www.fhi.ox.ac.uk/wp-content/uploads/2021/08/QNRs_FHI-TR-2021-3.0.pdf</code>`{=html}](https://www.fhi.ox.ac.uk/wp-content/uploads/2021/08/QNRs_FHI-TR-2021-3.0.pdf){.link-annotated-partial
    .id-not .include-annotation link-icon="pdf" link-icon-type="svg"
    link-icon-color="#f40f02"}**

-   **[`<code>https://www.lesswrong.com/posts/XaKLjyDejtXDoRAzL/a-quick-experiment-on-lms-inductive-biases-in-performing</code>`{=html}](https://www.lesswrong.com/posts/XaKLjyDejtXDoRAzL/a-quick-experiment-on-lms-inductive-biases-in-performing){.link-annotated-partial
    .id-not .include-annotation link-icon="LW" link-icon-type="text"
    link-icon-color="#7faf83"}**

-   **[`<code>https://www.lesswrong.com/posts/bwyKCQD7PFWKhELMr/by-default-gpts-think-in-plain-sight?commentId=zfzHshctWZYo8JkLe</code>`{=html}](https://www.lesswrong.com/posts/bwyKCQD7PFWKhELMr/by-default-gpts-think-in-plain-sight?commentId=zfzHshctWZYo8JkLe){.link-annotated-partial
    .id-not .include-annotation link-icon="LW" link-icon-type="text"
    link-icon-color="#7faf83"}**

-   **[`<code>https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-i/</code>`{=html}](https://www.oreilly.com/radar/what-we-learned-from-a-year-of-building-with-llms-part-i/){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://www.patterns.app/blog/2023/01/18/crunchbot-sql-analyst-gpt/</code>`{=html}](https://www.patterns.app/blog/2023/01/18/crunchbot-sql-analyst-gpt/){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://www.pnas.org/doi/full/10.1073/pnas.2317967121</code>`{=html}](https://www.pnas.org/doi/full/10.1073/pnas.2317967121){.link-annotated-partial
    .id-not .include-annotation link-icon="PNAS"
    link-icon-type="text,quad" link-icon-color="#1f75b9"}**

-   **[`<code>https://www.reddit.com/r/ChatGPT/comments/10zavbv/extending_chatgpt_with_some_additional_internal/</code>`{=html}](https://www.reddit.com/r/ChatGPT/comments/10zavbv/extending_chatgpt_with_some_additional_internal/){.link-annotated-partial
    .id-not .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.reddit.com/r/ChatGPT/comments/11anct1/its_easy_to_give_chatgpt_a_bonafide_consciousness/</code>`{=html}](https://www.reddit.com/r/ChatGPT/comments/11anct1/its_easy_to_give_chatgpt_a_bonafide_consciousness/){.link-annotated-partial
    .id-not .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.reddit.com/r/ChatGPT/comments/1pjitig/gemini_leaked_its_chain_of_thought_and_spiraled/</code>`{=html}](https://www.reddit.com/r/ChatGPT/comments/1pjitig/gemini_leaked_its_chain_of_thought_and_spiraled/){.id-not
    .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.reddit.com/r/LocalLLaMA/comments/1fuxw8d/just_for_kicks_i_looked_at_the_newly_released/</code>`{=html}](https://www.reddit.com/r/LocalLLaMA/comments/1fuxw8d/just_for_kicks_i_looked_at_the_newly_released/){.link-annotated-partial
    .id-not .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.reddit.com/r/OpenAI/comments/1fxa6d6/two_purported_instances_of_o1preview_and_o1mini/</code>`{=html}](https://www.reddit.com/r/OpenAI/comments/1fxa6d6/two_purported_instances_of_o1preview_and_o1mini/){.id-not
    .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.reddit.com/r/OpenAI/comments/1gjj430/o1_preview_got_weird_today/</code>`{=html}](https://www.reddit.com/r/OpenAI/comments/1gjj430/o1_preview_got_weird_today/){.id-not
    .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.reddit.com/r/PromptEngineering/comments/1fj6h13/hallucinations_in_o1preview_reasoning/</code>`{=html}](https://www.reddit.com/r/PromptEngineering/comments/1fj6h13/hallucinations_in_o1preview_reasoning/){.link-annotated-partial
    .id-not .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.reddit.com/r/slatestarcodex/comments/1201v68/10word_quote_a_short_and_simple_failure_mode_of/jdigzkh/?context=3</code>`{=html}](https://www.reddit.com/r/slatestarcodex/comments/1201v68/10word_quote_a_short_and_simple_failure_mode_of/jdigzkh/?context=3){.id-not
    .include-annotation link-icon="reddit" link-icon-type="svg"
    link-icon-color="#ff4500"}**

-   **[`<code>https://www.waluigipurple.com/post/revising-poetry-with-gpt-4</code>`{=html}](https://www.waluigipurple.com/post/revising-poetry-with-gpt-4){.link-annotated-partial
    .id-not .include-annotation}**

-   **[`<code>https://www.youtube.com/watch?v=g7YJIpkk7KM?t=38</code>`{=html}](https://www.youtube.com/watch?v=g7YJIpkk7KM?t=38){.link-annotated-partial
    .id-not .include-annotation link-icon="youtube" link-icon-type="svg"
    link-icon-color="#ff0033"}**

-   **[`<code>https://x.com/AISafetyMemes/status/1841891795782775221</code>`{=html}](https://x.com/AISafetyMemes/status/1841891795782775221 "AISafetyMemes 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/AmandaAskell/status/1986571451902927017</code>`{=html}](https://x.com/AmandaAskell/status/1986571451902927017 "Askell 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/AmandaAskell/status/1986571451902927017`](https://x.com/AmandaAskell/status/1986571451902927017 "Askell 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/BlinkDL_AI/status/1677593798531223552</code>`{=html}](https://x.com/BlinkDL_AI/status/1677593798531223552 "BlinkDL_AI 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/D_Rod_Tweets/status/1628449917898264576</code>`{=html}](https://x.com/D_Rod_Tweets/status/1628449917898264576 "D_Rod_Tweets 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/DaveMonlander/status/1612802240582135809</code>`{=html}](https://x.com/DaveMonlander/status/1612802240582135809 "DaveMonlander 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/Francis_YAO_/status/1884138762852262349</code>`{=html}](https://x.com/Francis_YAO_/status/1884138762852262349 "Yao 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/KevinAFischer/status/1646018246225846272</code>`{=html}](https://x.com/KevinAFischer/status/1646018246225846272 "KevinAFischer 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/KevinAFischer/status/1646677902833102849</code>`{=html}](https://x.com/KevinAFischer/status/1646677902833102849 "KevinAFischer 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/KevinAFischer/status/1646690838981005312</code>`{=html}](https://x.com/KevinAFischer/status/1646690838981005312 "KevinAFischer 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/Kyrannio/status/1793874431179460911</code>`{=html}](https://x.com/Kyrannio/status/1793874431179460911 "Kyrannio 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/MParakhin/status/1632087709060825088</code>`{=html}](https://x.com/MParakhin/status/1632087709060825088 "Parakhin 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/MikePFrank/status/1622202768743096320</code>`{=html}](https://x.com/MikePFrank/status/1622202768743096320 "MikePFrank 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/MikePFrank/status/1622495004810784768</code>`{=html}](https://x.com/MikePFrank/status/1622495004810784768 "MikePFrank 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/Shalev_lif/status/1883214799938396356</code>`{=html}](https://x.com/Shalev_lif/status/1883214799938396356 "Shalev_lif 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/StudentInfosec/status/1640360234882310145</code>`{=html}](https://x.com/StudentInfosec/status/1640360234882310145 "StudentInfosec 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/Tim_Hua_/status/2044300361348067457</code>`{=html}](https://x.com/Tim_Hua_/status/2044300361348067457 "Hua 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/adamhfry/status/1972340656124387491</code>`{=html}](https://x.com/adamhfry/status/1972340656124387491 "adamhfry 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/adiwyner/status/1629980541716922369</code>`{=html}](https://x.com/adiwyner/status/1629980541716922369 "adiwyner 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/aidenybai/status/1993901129210712129</code>`{=html}](https://x.com/aidenybai/status/1993901129210712129 "aidenybai 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/amasad/status/1628546489843863555</code>`{=html}](https://x.com/amasad/status/1628546489843863555 "Masad 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/anderssandberg/status/1881839255028474306</code>`{=html}](https://x.com/anderssandberg/status/1881839255028474306 "anderssandberg 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/andrewwhite01/status/1616933106786738176</code>`{=html}](https://x.com/andrewwhite01/status/1616933106786738176 "White 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/chehendriksen/status/1895873675767005254</code>`{=html}](https://x.com/chehendriksen/status/1895873675767005254 "chehendriksen 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/chehendriksen/status/1895873675767005254`](https://x.com/chehendriksen/status/1895873675767005254 "chehendriksen 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/chrisbarber/status/1885047105741611507</code>`{=html}](https://x.com/chrisbarber/status/1885047105741611507 "Barber 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/chrisbarber/status/1885047105741611507`](https://x.com/chrisbarber/status/1885047105741611507 "Barber 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/deepfates/status/1682110624271319040</code>`{=html}](https://x.com/deepfates/status/1682110624271319040 "deepfates 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/denny_zhou/status/1547662872511070212</code>`{=html}](https://x.com/denny_zhou/status/1547662872511070212 "denny_zhou 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/denny_zhou/status/1587115933293678592</code>`{=html}](https://x.com/denny_zhou/status/1587115933293678592 "denny_zhou 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/emollick/status/1705422957856604503</code>`{=html}](https://x.com/emollick/status/1705422957856604503 "Mollick 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/emollick/status/2064542441848422611</code>`{=html}](https://x.com/emollick/status/2064542441848422611 "Mollick 2026"){.link-modified-recently
    .link-annotated-partial .id-not .include-annotation
    link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/finereli/status/1782611247709786145</code>`{=html}](https://x.com/finereli/status/1782611247709786145 "finereli 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/finereli/status/1782611247709786145`](https://x.com/finereli/status/1782611247709786145 "finereli 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/gfodor/status/1626270272314839041</code>`{=html}](https://x.com/gfodor/status/1626270272314839041 "gfodor 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1563191853587271681</code>`{=html}](https://x.com/goodside/status/1563191853587271681 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1568375796904886274</code>`{=html}](https://x.com/goodside/status/1568375796904886274 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1568375802903015425</code>`{=html}](https://x.com/goodside/status/1568375802903015425 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1568416130133368835</code>`{=html}](https://x.com/goodside/status/1568416130133368835 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1568448128495534081</code>`{=html}](https://x.com/goodside/status/1568448128495534081 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1581868987952300032</code>`{=html}](https://x.com/goodside/status/1581868987952300032 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1612017392518840320</code>`{=html}](https://x.com/goodside/status/1612017392518840320 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/goodside/status/1635711013566795776</code>`{=html}](https://x.com/goodside/status/1635711013566795776 "Goodside 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/jconorgrogan/status/2073443593268650212</code>`{=html}](https://x.com/jconorgrogan/status/2073443593268650212 "jconorgrogan 2026"){.link-modified-recently
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/jconorgrogan/status/2073443593268650212`](https://x.com/jconorgrogan/status/2073443593268650212 "jconorgrogan 2026"){.link-modified-recently
    .id-not .include-content}
    :::

-   **[`<code>https://x.com/jd_pressman/status/1646766004637401088</code>`{=html}](https://x.com/jd_pressman/status/1646766004637401088 "jd_pressman 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/jeremyphoward/status/1801037736968913128</code>`{=html}](https://x.com/jeremyphoward/status/1801037736968913128 "Howard 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/jeremyphoward/status/1801037736968913128`](https://x.com/jeremyphoward/status/1801037736968913128 "Howard 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/jmilldotdev/status/1592288240861839360</code>`{=html}](https://x.com/jmilldotdev/status/1592288240861839360 "jmilldotdev 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/lemonodor/status/1628270074074398720</code>`{=html}](https://x.com/lemonodor/status/1628270074074398720 "lemonodor 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/littmath/status/1598128056874721283</code>`{=html}](https://x.com/littmath/status/1598128056874721283 "littmath 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/mbusigin/status/1789334007047455178</code>`{=html}](https://x.com/mbusigin/status/1789334007047455178 "mbusigin 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/mbusigin/status/1789334007047455178`](https://x.com/mbusigin/status/1789334007047455178 "mbusigin 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/md_rumpf/status/1647911393796956162</code>`{=html}](https://x.com/md_rumpf/status/1647911393796956162 "md_rumpf 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/peterwildeford/status/1522633978305560576</code>`{=html}](https://x.com/peterwildeford/status/1522633978305560576 "peterwildeford 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/polynoamial/status/1903501780102926552</code>`{=html}](https://x.com/polynoamial/status/1903501780102926552 "Brown 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/polynoamial/status/1903501780102926552`](https://x.com/polynoamial/status/1903501780102926552 "Brown 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/repligate/status/1884791264035602476</code>`{=html}](https://x.com/repligate/status/1884791264035602476 "Janus 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/sand7one/status/1912012664269930896</code>`{=html}](https://x.com/sand7one/status/1912012664269930896 "sand7one 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/shinboson/status/1805459742518595585</code>`{=html}](https://x.com/shinboson/status/1805459742518595585 "shinboson 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://x.com/wgussml/status/1834712489822765295</code>`{=html}](https://x.com/wgussml/status/1834712489822765295 "wgussml 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/wgussml/status/1834712489822765295`](https://x.com/wgussml/status/1834712489822765295 "wgussml 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/willccbb/status/1885927147505647832</code>`{=html}](https://x.com/willccbb/status/1885927147505647832 "willccbb 2026"){.id-not
    .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

    ::: aux-links-transclude-file
    [`https://x.com/willccbb/status/1885927147505647832`](https://x.com/willccbb/status/1885927147505647832 "willccbb 2026"){.id-not
    .include-content}
    :::

-   **[`<code>https://x.com/yoheinakajima/status/1670557048743010305</code>`{=html}](https://x.com/yoheinakajima/status/1670557048743010305 "yoheinakajima 2026"){.link-annotated-partial
    .id-not .include-annotation link-icon="twitter" link-icon-type="svg"
    link-icon-color="#1da1f2"}**

-   **[`<code>https://yaofu.notion.site/A-Closer-Look-at-Large-Language-Models-Emergent-Abilities-493876b55df5479d80686f68a1abd72f</code>`{=html}](https://yaofu.notion.site/A-Closer-Look-at-Large-Language-Models-Emergent-Abilities-493876b55df5479d80686f68a1abd72f){.link-annotated-partial
    .id-not .include-annotation}**

# Bibliography {#link-bibliography-section}

1.  `https://www.lowimpactfruit.com/p/zork-bench-an-llm-reasoning-eval`:
    [`“Zork-Bench: An LLM Reasoning Eval Based on Text Adventure Games; a Tale As Old As Time, or at Least As Old As Computers”`{=html}](https://www.lowimpactfruit.com/p/zork-bench-an-llm-reasoning-eval "'zork-bench: An LLM reasoning eval based on text adventure games; a tale as old as time, or at least as old as computers', Aiken 2026"){#aiken-2026
    .link-annotated}, [John Aiken]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fwww.lowimpactfruit.com%252Fp%252Fzork-bench-an-llm-reasoning-eval.html "Directory-tag link-bibliography for link https://www.lowimpactfruit.com/p/zork-bench-an-llm-reasoning-eval"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

2.  `https://arxiv.org/abs/2512.02556#deepseek`:
    [`“DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models”`{=html}](https://arxiv.org/abs/2512.02556#deepseek "'DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models', DeepSeek et al 2025"){#deepseek-et-al-2025
    .link-annotated},
    [[DeepSeek](https://www.deepseek.com/ "DeepSeek"){.id-not}, Aixin Liu, Aoxue Mei[, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, [Hao Li](https://en.wikipedia.org/wiki/Hao_Li "Hao Li"){.id-not}, Haofen Liang, Haoran Wei, Haowei Zhang, Haowen Luo, Haozhe Ji, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, Jialiang Huang, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jingchang Chen, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jinhua Zhu, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexin Huang, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Liang Zhao, Liangsheng Yin, Lihua Guo, Lingxiao Luo, Linwang Ma, Litong Wang, Liyue Zhang, M. S. Di, M. Y. Xu, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Panpan Huang, Peixin Cong, Peiyi Wang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, S. H. Liu, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Songyang Zhou, Tao Ni, Tao Yun, Tian Pei, Tian Ye, Tianyuan Yue, Wangding Zeng, Wen Liu, [Liang Wenfeng](https://en.wikipedia.org/wiki/Liang_Wenfeng "Liang Wenfeng"){.id-not}, Wenjie Pang, Wenjing Luo, Wenjun Gao, Wentao Zhang, Xi Gao, Xiangwen Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyuan Li, [Xu Chen](https://en.wikipedia.org/wiki/Xu_Chen "Xu Chen"){.id-not}, Xuecheng Su, Xuehai Pan, Xuheng Lin, Xuwei Fu, Y. Q. Wang, Yang Zhang, Yanhong Xu, Yanru Ma, Yao Li, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, [Yi Yu](https://en.wikipedia.org/wiki/Yi_Yu "Yi Yu"){.id-not}, Yichao Zhang, Yifan Ding, Yifan Shi, Yiliang Xiong, Ying He, Ying Zhou, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuduan Wang, Yue Gong, Yuhan Wu, Yuheng Zou, Yukun Li, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Z. F. Wu, Z. Z. Ren, Zehua Zhao, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhiyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Zizheng Pan, Zongqing Yao, Bei Feng, Hui Li, J. L. Cai, Jiaqi Ni, [Lei Xu](https://en.wikipedia.org/wiki/Lei_Xu "Lei Xu"){.id-not}, Meng Li, Ning Tian, R. J. Chen, R. L. Jin, S. S. Li, Shuang Zhou, Tianyu Sun, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xinnan Song, Xinyi Zhou, Y. X. Zhu, [Yanping Huang](https://scholar.google.com/citations?user=uEtBQScAAAAJ "Yanping Huang"){.id-not}, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, Dongjie Ji, Jian Liang, Jianzhong Guo, Jin Chen, Leyi Xia, Miaojun Wang, Mingming Li, [Peng Zhang](https://sites.google.com/site/pengzhang27182/ "Peng Zhang’s homepage"){.id-not}, Ruyi Chen, Shangmian Sun, Shaoqing Wu, Shengfeng Ye, T. Wang, W. L. Xiao, Wei An, Xianzu Wang, Xiaowen Sun, Xiaoxiang Wang, Ying Tang, Yukun Zha, Zekai Zhang, Zhe Ju, Zhen Zhang, Zihua Qu]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2512.02556%2523deepseek.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2512.02556#deepseek"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

3.  `https://openai.com/index/introducing-gpt-5/#gpt-5-pro`:
    [`“GPT-5 Pro: Scaled but Efficient Parallel Test-Time Compute, to Provide the Highest Quality and Most Comprehensive Answers”`{=html}](https://openai.com/index/introducing-gpt-5/#gpt-5-pro "'GPT-5 Pro: scaled but efficient parallel test-time compute, to provide the highest quality and most comprehensive answers', OpenAI 2025"){#openai-2025
    .link-annotated},
    [[OpenAI](https://en.wikipedia.org/wiki/OpenAI "OpenAI"){.link-annotated-partial
    .id-not}]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fopenai.com%252Findex%252Fintroducing-gpt-5%252F%2523gpt-5-pro.html "Directory-tag link-bibliography for link https://openai.com/index/introducing-gpt-5/#gpt-5-pro"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

4.  `https://arxiv.org/abs/2506.10922`:
    [`“Robustly Improving LLM Fairness in Realistic Settings via Interpretability”`{=html}](https://arxiv.org/abs/2506.10922 "'Robustly Improving LLM Fairness in Realistic Settings via Interpretability', Karvonen & Marks 2025"){#karvonen-marks-2025
    .link-annotated},
    [Adam Karvonen, [Samuel Marks](https://scholar.google.com/citations?user=fW7yK10AAAAJ "Samuel Marks"){.id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2506.10922.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2506.10922"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

5.  `https://arxiv.org/abs/2505.11711`:
    [`“Reinforcement Learning Finetunes Small Subnetworks in Large Language Models”`{=html}](https://arxiv.org/abs/2505.11711 "'Reinforcement Learning Finetunes Small Subnetworks in Large Language Models', Mukherjee et al 2025"){#mukherjee-et-al-2025
    .link-annotated},
    [Sagnik Mukherjee, [Lifan Yuan](https://en.wikipedia.org/wiki/Lifan_Yuan "Lifan Yuan"){.id-not}, Dilek Hakkani-Tur, Hao Peng]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2505.11711.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2505.11711"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

6.  `https://arxiv.org/abs/2504.20571`:
    [`“Reinforcement Learning for Reasoning in Large Language Models With One Training Example”`{=html}](https://arxiv.org/abs/2504.20571 "'Reinforcement Learning for Reasoning in Large Language Models with One Training Example', Wang et al 2025"){#wang-et-al-2025
    .link-annotated},
    [Yiping Wang, Qing Yang, Zhiyuan Zeng[, Liliang Ren, Lucas Liu, Baolin Peng, Hao Cheng, Xuehai He, Kuan Wang, [Jianfeng Gao](https://www.microsoft.com/en-us/research/people/jfgao/ "Jianfeng Gao at Microsoft Research"){.link-annotated-partial
    .id-not}, [Weizhu Chen](https://scholar.google.com/citations?user=LG_E-4EAAAAJ "Weizhu Chen"){.id-not}, Shuohang Wang, Simon Shaolei Du, Yelong Shen]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2504.20571.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2504.20571"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

7.  `https://arxiv.org/abs/2504.14379`:
    [`“The Geometry of Self-Verification in a Task-Specific Reasoning Model”`{=html}](https://arxiv.org/abs/2504.14379 "'The Geometry of Self-Verification in a Task-Specific Reasoning Model', Lee et al 2025"){#lee-et-al-2025
    .link-annotated},
    [Andrew Lee, Lihao Sun, Chris Wendler[, [Fernanda Viégas](https://en.wikipedia.org/wiki/Fernanda_Vi%C3%A9gas "Fernanda Viégas"){.id-not}, [Martin M. Wattenberg](https://en.wikipedia.org/wiki/Martin_M._Wattenberg "Martin M. Wattenberg"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2504.14379.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2504.14379"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

8.  `https://arxiv.org/abs/2503.16219`:
    [`“Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t”`{=html}](https://arxiv.org/abs/2503.16219 "'Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn’t', Dang & Ngo 2025"){#dang-ngo-2025
    .link-annotated}, [Quy-Anh Dang, Chris Ngo]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2503.16219.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2503.16219"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

9.  `https://arxiv.org/abs/2502.20339`:
    [`“Thinking Slow, Fast: Scaling Inference Compute With Distilled Reasoners”`{=html}](https://arxiv.org/abs/2502.20339 "'Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners', Paliotta et al 2025"){#paliotta-et-al-2025
    .link-annotated},
    [Daniele Paliotta, [Junxiong Wang](https://www.cs.cornell.edu/~junxiong/ "Junxiong Wang"){.link-annotated-partial
    .id-not}, Matteo Pagliardini[, Kevin Y. Li, Aviv Bick, [J. Zico Kolter](https://en.wikipedia.org/wiki/Zico_Kolter "Zico Kolter"){.id-not}, [Albert Gu](https://scholar.google.com/citations?user=DVCHv1kAAAAJ "Albert Gu"){.id-not}, François Fleuret, [Tri Dao](https://tridao.me/ "Tri Dao"){.link-annotated-partial
    .id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2502.20339.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2502.20339"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

10. `https://arxiv.org/abs/2502.06807#openai`:
    [`“Competitive Programming With Large Reasoning Models”`{=html}](https://arxiv.org/abs/2502.06807#openai "'Competitive Programming with Large Reasoning Models', El-Kishky et al 2025"){#el-kishky-et-al-2025
    .link-annotated},
    [Ahmed El-Kishky, Alexander Wei, Andre Saraiva[, Borys Minaev, [Daniel Selsam](https://scholar.google.com/citations?user=yaSqFaEAAAAJ "Daniel Selsam"){.id-not}, [David Dohan](https://scholar.google.com/citations?user=iZ5cY0AAAAAJ&view_op=list_works&sortby=pubdate "David Dohan"){.id-not}, Francis Song, Hunter Lightman, Ignasi Clavera, [Jakub Pachocki](https://en.wikipedia.org/wiki/Jakub_Pachocki "Jakub Pachocki"){.link-modified-recently
    .id-not}, Jerry Tworek, Lorenz Kuhn, [Łukasz Kaiser](https://scholar.google.com/citations?user=JWmiQR0AAAAJ "Łukasz Kaiser"){.id-not}, [Mark Chen](https://event.technologyreview.com/emtech-mit-2023/speaker/901826/mark-chen "Speaker Details: EmTech MIT 2023"){.link-annotated-partial
    .id-not}, Max Schwarzer, Mostafa Rohaninejad, [Nat McAleese](https://scholar.google.com/citations?user=crw6TeIAAAAJ "Nat McAleese"){.id-not}, o3 contributors, Oleg Mürk, Rhythm Garg, Rui Shu, Szymon Sidor, Vineet Kosaraju, Wenda Zhou]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2502.06807%2523openai.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2502.06807#openai"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

11. `https://arxiv.org/abs/2501.08156`:
    [`“Are DeepSeek R1 And Other Reasoning Models More Faithful?”`{=html}](https://arxiv.org/abs/2501.08156 "'Are DeepSeek R1 And Other Reasoning Models More Faithful?', Chua & Evans 2025"){#chua-evans-2025
    .link-annotated},
    [James Chua, [Owain Evans](https://owainevans.github.io/ "Owain Evans, AI Alignment researcher"){.link-annotated-partial
    .id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2501.08156.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2501.08156"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

12. `https://arxiv.org/abs/2412.01981`:
    [`“Free Process Rewards without Process Labels”`{=html}](https://arxiv.org/abs/2412.01981 "'Free Process Rewards without Process Labels', Yuan et al 2024"){#yuan-et-al-2024
    .link-annotated},
    [[Lifan Yuan](https://en.wikipedia.org/wiki/Lifan_Yuan "Lifan Yuan"){.id-not}, Wendi Li, Huayu Chen[, Ganqu Cui, [Ning Ding](https://scholar.google.com/citations?user=uZXQuYAAAAAJ "Ning Ding"){.id-not}, Kaiyan Zhang, Bowen Zhou, [Zhiyuan Liu](https://nlp.csai.tsinghua.edu.cn/~lzy/ "Zhiyuan Liu"){.link-annotated-partial
    .id-not}, Hao Peng]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2412.01981.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2412.01981"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

13. `https://arxiv.org/abs/2410.21333`:
    [`“Mind Your Step (By Step): Chain-Of-Thought Can Reduce Performance on Tasks Where Thinking Makes Humans Worse”`{=html}](https://arxiv.org/abs/2410.21333 "'Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse', Liu et al 2024"){#liu-et-al-2024
    .link-annotated},
    [Ryan Liu, Jiayi Geng, Addison J. Wu[, Ilia Sucholutsky, Tania Lombrozo, [Thomas L. Griffiths](https://en.wikipedia.org/wiki/Thomas_L._Griffiths "Thomas L. Griffiths"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2410.21333.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2410.21333"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

14. `https://arxiv.org/abs/2406.13121#google`:
    [`“Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?”`{=html}](https://arxiv.org/abs/2406.13121#google "'Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?', Lee et al 2024"){#lee-et-al-2024-1
    .link-annotated},
    [Jinhyuk Lee, [Anthony Chen](https://en.wikipedia.org/wiki/Anthony_Chen "Anthony Chen"){.id-not}, Zhuyun Dai[, Dheeru Dua, Devendra Singh Sachan, Michael Boratko, Yi Luan, Sébastien M. R. Arnold, Vincent Perot, Siddharth Dalmia, Hexiang Hu, Xudong Lin, Panupong Pasupat, Aida Amini, Jeremy R. Cole, [Sebastian Riedel](http://www.riedelcastro.org/ "'Sebastian Riedel homepage', Riedel 2026"){.link-annotated-partial
    .id-not}, Iftekhar Naim, Ming-Wei Chang, Kelvin Guu]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2406.13121%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2406.13121#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

15. `https://arxiv.org/abs/2405.15143`:
    [`“Intelligent Go-Explore (IGE): Standing on the Shoulders of Giant Foundation Models”`{=html}](https://arxiv.org/abs/2405.15143 "'Intelligent Go-Explore (IGE): Standing on the Shoulders of Giant Foundation Models', Lu et al 2024"){#lu-et-al-2024-2
    .link-annotated},
    [Cong Lu, Shengran Hu, [Jeff Clune](http://jeffclune.com/ "Jeff Clune—Professor—Computer Science—University of British Columbia"){.link-annotated-partial
    .id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2405.15143.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2405.15143"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

16. `https://arxiv.org/abs/2405.14838`:
    [`“From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step”`{=html}](https://arxiv.org/abs/2405.14838 "'From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step', Deng et al 2024"){#deng-et-al-2024-1
    .link-annotated},
    [Yuntian Deng, [Yejin Choi](https://en.wikipedia.org/wiki/Yejin_Choi "Yejin Choi"){.id-not}, Stuart Shieber]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2405.14838.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2405.14838"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

17. `https://arxiv.org/abs/2405.10938`:
    [`“Observational Scaling Laws and the Predictability of Language Model Performance”`{=html}](https://arxiv.org/abs/2405.10938 "'Observational Scaling Laws and the Predictability of Language Model Performance', Ruan et al 2024"){#ruan-et-al-2024
    .link-annotated},
    [Yangjun Ruan, Chris J. Maddison, [Tatsunori Hashimoto](https://thashim.github.io/ "Tatsunori Hashimoto"){.link-annotated-partial
    .id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2405.10938.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2405.10938"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

18. `https://arxiv.org/abs/2404.15574`:
    [`“Retrieval Head Mechanistically Explains Long-Context Factuality”`{=html}](https://arxiv.org/abs/2404.15574 "'Retrieval Head Mechanistically Explains Long-Context Factuality', Wu et al 2024"){#wu-et-al-2024-1
    .link-annotated},
    [Wenhao Wu, [Yizhong Wang](https://homes.cs.washington.edu/~yizhongw/ "Yizhong Wang—University of Washington"){.link-annotated-partial
    .id-not}, Guangxuan Xiao[, Hao Peng, Yao Fu]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2404.15574.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2404.15574"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

19. `https://arxiv.org/abs/2404.15758`:
    [`“Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models”`{=html}](https://arxiv.org/abs/2404.15758 "'Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models', Pfau et al 2024"){#pfau-et-al-2024
    .link-annotated},
    [Jacob Pfau, William Merrill, [Samuel R. Bowman](https://cims.nyu.edu/~sbowman/ "Sam Bowman"){.link-annotated-partial
    .id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2404.15758.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2404.15758"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

20. `https://link.springer.com/article/10.1007/s10506-024-09396-9`:
    [`“Re-Evaluating GPT-4’s Bar Exam Performance”`{=html}](https://link.springer.com/article/10.1007/s10506-024-09396-9 "'Re-evaluating GPT-4’s bar exam performance', Martínez 2024"){#martínez-2024
    .link-annotated}, [Eric Martínez]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Flink.springer.com%252Farticle%252F10.1007%252Fs10506-024-09396-9.html "Directory-tag link-bibliography for link https://link.springer.com/article/10.1007/s10506-024-09396-9"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

21. `https://arxiv.org/abs/2403.18802#deepmind`:
    [`“Long-Form Factuality in Large Language Models”`{=html}](https://arxiv.org/abs/2403.18802#deepmind "'Long-form factuality in large language models', Wei et al 2024"){#wei-et-al-2024-1
    .link-annotated},
    [Jerry Wei, Chengrun Yang, Xinying Song[, Yifeng Lu, Nathan Hu, Jie Huang, Dustin Tran, Daiyi Peng, Ruibo Liu, Da Huang, Cosmo Du, [Quoc V. Le](https://en.wikipedia.org/wiki/Quoc_V._Le "Quoc V. Le"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2403.18802%2523deepmind.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2403.18802#deepmind"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

22. `https://arxiv.org/abs/2403.18120#google`:
    [`“Don’t Trust: Verify—Grounding LLM Quantitative Reasoning With Autoformalization”`{=html}](https://arxiv.org/abs/2403.18120#google "'Don’t Trust: Verify—Grounding LLM Quantitative Reasoning with Autoformalization', Zhou et al 2024"){#zhou-et-al-2024
    .link-annotated},
    [Jin Peng Zhou, Charles Staats, Wenda Li[, Christian Szegedy, [Kilian Q. Weinberger](https://www.cs.cornell.edu/~kilian/ "Welcome"){.link-annotated-partial
    .id-not}, [Yuhuai Wu](https://yuhuaiwu.github.io/ "Yuhuai (Tony) Wu’s Home Page"){.link-annotated-partial
    .id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2403.18120%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2403.18120#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

23. `https://arxiv.org/abs/2403.09629`:
    [`“Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking”`{=html}](https://arxiv.org/abs/2403.09629 "'Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking', Zelikman et al 2024"){#zelikman-et-al-2024
    .link-annotated},
    [Eric Zelikman, Georges Harik, Yijia Shao[, Varuna Jayasiri, Nick Haber, [Noah D. Goodman](https://scholar.google.com/citations?user=OUpIbcQAAAAJ "Noah D. Goodman"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2403.09629.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2403.09629"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

24. `https://arxiv.org/abs/2402.14903`:
    [`“Tokenization Counts: the Impact of Tokenization on Arithmetic in Frontier LLMs”`{=html}](https://arxiv.org/abs/2402.14903 "'Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs', Singh & Strouse 2024"){#singh-strouse-2024
    .link-annotated}, [Aaditya K. Singh, D. J. Strouse]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2402.14903.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2402.14903"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

25. `https://arxiv.org/abs/2402.09963`:
    [`“Why Are Sensitive Functions Hard for Transformers?”`{=html}](https://arxiv.org/abs/2402.09963 "'Why are Sensitive Functions Hard for Transformers?', Hahn & Rofin 2024"){#hahn-rofin-2024
    .link-annotated},
    [[Michael Hahn](https://en.wikipedia.org/wiki/Michael_Hahn "Michael Hahn"){.id-not}, Mark Rofin]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2402.09963.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2402.09963"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

26. `https://arxiv.org/abs/2402.05120#tencent`:
    [`“More Agents Is All You Need”`{=html}](https://arxiv.org/abs/2402.05120#tencent "'More Agents Is All You Need', Li et al 2024"){#li-et-al-2024-10
    .link-annotated},
    [Junyou Li, [Qin Zhang](https://en.wikipedia.org/wiki/Qin_Zhang "Qin Zhang"){.id-not}, Yangbin Yu[, Qiang Fu, Deheng Ye]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2402.05120%2523tencent.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2402.05120#tencent"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

27. `https://arxiv.org/abs/2312.08935`:
    [`“Math-Shepherd: Verify and Reinforce LLMs Step-By-Step without Human Annotations”`{=html}](https://arxiv.org/abs/2312.08935 "'Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations', Wang et al 2023"){#wang-et-al-2023
    .link-annotated},
    [Peiyi Wang, Lei Li, Zhihong Shao[, R. X. Xu, Damai Dai, Yifei Li, Deli Chen, Y. Wu, Zhifang Sui]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2312.08935.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2312.08935"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

28. `https://arxiv.org/abs/2312.06585#deepmind`:
    [`“Beyond Human Data: Scaling Self-Training for Problem-Solving With Language Models (ReST<sup>EM</sup>)”`{=html}](https://arxiv.org/abs/2312.06585#deepmind "'Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models (ReST<sup>EM</sup>)', Singh et al 2023"){#singh-et-al-2023-3
    .link-annotated},
    [Avi Singh, John D. Co-Reyes, Rishabh Agarwal[, Ankesh Anand, Piyush Patil, Peter J. Liu, James Harrison, Jaehoon Lee, Kelvin Xu, Aaron Parisi, [Abhishek Kumar](https://scholar.google.com/citations?user=6vghMS0AAAAJ "Abhishek Kumar"){.id-not}, Alex Alemi, Alex Rizkowsky, Azade Nova, Ben Adlam, Bernd Bohnet, Hanie Sedghi, [Igor Mordatch](https://scholar.google.com/citations?user=Vzr1RukAAAAJ "Igor Mordatch"){.id-not}, Isabelle Simpson, Izzeddin Gur, Jasper Snoek, [Jeffrey Pennington](https://en.wikipedia.org/wiki/Jeffrey_Pennington "Jeffrey Pennington"){.id-not}, Jiri Hron, [Kathleen Kenealy](https://en.wikipedia.org/wiki/Kathleen_Kenealy "Kathleen Kenealy"){.id-not}, Kevin Swersky, Kshiteej Mahajan, Laura Culp, Lechao Xiao, Maxwell L. Bileschi, [Noah Constant](https://scholar.google.com/citations?user=PbgcS6AAAAAJ "Noah Constant"){.id-not}, Roman Novak, [Rosanne Liu](https://rosanneliu.com/ "Rosanne Liu"){.link-annotated-partial
    .id-not}, Tris Warkentin, Yundi Qian, [Ethan Dyer](https://scholar.google.com/citations?user=LWeVRdUAAAAJ "Ethan Dyer"){.id-not}, [Behnam Neyshabur](https://www.neyshabur.net/ "Behnam Neyshabur"){.id-not}, [Jascha Sohl-Dickstein](https://scholar.google.com/citations?user=-3zYIjQAAAAJ "Jascha Sohl-Dickstein"){.id-not}, Noah Fiedel]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2312.06585%2523deepmind.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2312.06585#deepmind"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

29. `https://arxiv.org/abs/2311.16452#microsoft`:
    [`“Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine”`{=html}](https://arxiv.org/abs/2311.16452#microsoft "'Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine', Nori et al 2023"){#nori-et-al-2023
    .link-annotated},
    [Harsha Nori, Yin Tat Lee, Sheng Zhang[, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, Renqian Luo, Scott Mayer McKinney, Robert Osazuwa Ness, Hoifung Poon, [Tao Qin](https://en.wikipedia.org/wiki/Tao_Qin "Tao Qin"){.id-not}, Naoto Usuyama, Chris White, [Eric Horvitz](https://en.wikipedia.org/wiki/Eric_Horvitz "Eric Horvitz"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2311.16452%2523microsoft.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2311.16452#microsoft"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

30. `https://arxiv.org/abs/2312.02179`:
    [`“Training Chain-Of-Thought via Latent-Variable Inference”`{=html}](https://arxiv.org/abs/2312.02179 "'Training Chain-of-Thought via Latent-Variable Inference', Phan et al 2023"){#phan-et-al-2023
    .link-annotated},
    [Du Phan, Matthew D. Hoffman, [David Dohan](https://scholar.google.com/citations?user=iZ5cY0AAAAAJ&view_op=list_works&sortby=pubdate "David Dohan"){.id-not}[, [Sholto Douglas](https://en.wikipedia.org/wiki/Sholto_Douglas "Sholto Douglas"){.id-not}, Tuan Anh Le, Aaron Parisi, Pavel Sountsov, Charles Sutton, Sharad Vikram, Rif A. Saurous]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2312.02179.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2312.02179"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

31. `https://arxiv.org/abs/2310.08678`:
    [`“Can GPT Models Be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on Mock CFA Exams”`{=html}](https://arxiv.org/abs/2310.08678 "'Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams', Callanan et al 2023"){#callanan-et-al-2023
    .link-annotated},
    [Ethan Callanan, Amarachi Mbakwe, Antony Papadimitriou[, Yulong Pei, Mathieu Sibue, Xiaodan Zhu, Zhiqiang Ma, Xiaomo Liu, Sameena Shah]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2310.08678.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2310.08678"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

32. `https://arxiv.org/abs/2310.04406`:
    [`“Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models”`{=html}](https://arxiv.org/abs/2310.04406 "'Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models', Zhou et al 2023"){#zhou-et-al-2023-04
    .link-annotated},
    [Andy Zhou, Kai Yan, Michal Shlapentokh-Rothman[, Haohan Wang, Yu-Xiong Wang]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2310.04406.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2310.04406"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

33. `https://arxiv.org/abs/2310.02226`:
    [`“Think Before You Speak: Training Language Models With Pause Tokens”`{=html}](https://arxiv.org/abs/2310.02226 "'Think before you speak: Training Language Models With Pause Tokens', Goyal et al 2023"){#goyal-et-al-2023
    .link-annotated},
    [Sachin Goyal, Ziwei Ji, Ankit Singh Rawat[, Aditya Krishna Menon, [Sanjiv Kumar](https://scholar.google.com/citations?user=08CNqrYAAAAJ&view_op=list_works&sortby=pubdate "Sanjiv Kumar"){.id-not}, Vaishnavh Nagarajan]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2310.02226.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2310.02226"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

34. `https://arxiv.org/abs/2309.09117#facebook`:
    [`“Contrastive Decoding Improves Reasoning in Large Language Models”`{=html}](https://arxiv.org/abs/2309.09117#facebook "'Contrastive Decoding Improves Reasoning in Large Language Models', O’Brien & Lewis 2023"){#obrien-lewis-2023
    .link-annotated},
    [Sean O'Brien, [Mike Lewis](https://scholar.google.com/citations?user=SnQnQicAAAAJ "Mike Lewis"){.id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2309.09117%2523facebook.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2309.09117#facebook"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

35. `https://arxiv.org/abs/2309.06275`:
    [`“Re2: Re-Reading Improves Reasoning in Large Language Models”`{=html}](https://arxiv.org/abs/2309.06275 "'Re2: Re-Reading Improves Reasoning in Large Language Models', Xu et al 2023"){#xu-et-al-2023
    .link-annotated},
    [Xiaohan Xu, Chongyang Tao, Tao Shen[, [Can Xu](https://en.wikipedia.org/wiki/Can_Xu "Can Xu"){.id-not}, Hongbo Xu, Guodong Long, Jian-guang Lou, Shuai Ma]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2309.06275.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2309.06275"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

36. `https://arxiv.org/abs/2309.04269`:
    [`“From Sparse to Dense: GPT-4 Summarization With Chain of Density (CoD) Prompting”`{=html}](https://arxiv.org/abs/2309.04269 "'From Sparse to Dense: GPT-4 Summarization with Chain of Density (CoD) Prompting', Adams et al 2023"){#adams-et-al-2023
    .link-annotated},
    [Griffin Adams, Alexander Fabbri, Faisal Ladhak[, Eric Lehman, [Noémie Elhadad](https://en.wikipedia.org/wiki/No%C3%A9mie_Elhadad "Noémie Elhadad"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2309.04269.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2309.04269"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

37. `https://arxiv.org/abs/2308.07921`:
    [`“Solving Challenging Math Word Problems Using GPT-4 Code Interpreter With Code-Based Self-Verification”`{=html}](https://arxiv.org/abs/2308.07921 "'Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification', Zhou et al 2023"){#zhou-et-al-2023-06
    .link-annotated},
    [Aojun Zhou, [Ke Wang](https://en.wikipedia.org/wiki/Ke_Wang "Ke Wang"){.id-not}, Zimu Lu[, Weikang Shi, Sichun Luo, Zipeng Qin, Shaoqing Lu, Anya Jia, Linqi Song, Mingjie Zhan, [Hongsheng Li](https://www.ee.cuhk.edu.hk/~hsli/){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2308.07921.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2308.07921"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

38. `https://arxiv.org/abs/2307.05300#microsoft`:
    [`“Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration”`{=html}](https://arxiv.org/abs/2307.05300#microsoft "'Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration', Wang et al 2023"){#wang-et-al-2023-12
    .link-annotated},
    [Zhenhailong Wang, Shaoguang Mao, Wenshan Wu[, Tao Ge, [Furu Wei](https://scholar.google.com/citations?user=G-V1VpwAAAAJ "Furu Wei"){.id-not}, [Heng Ji](https://en.wikipedia.org/wiki/Heng_Ji "Heng Ji"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2307.05300%2523microsoft.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2307.05300#microsoft"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

39. `https://arxiv.org/abs/2307.03381`:
    [`“Teaching Arithmetic to Small Transformers”`{=html}](https://arxiv.org/abs/2307.03381 "'Teaching Arithmetic to Small Transformers', Lee et al 2023"){#lee-et-al-2023-2
    .link-annotated},
    [Nayoung Lee, Kartik Sreenivasan, Jason D. Lee[, Kangwook Lee, Dimitris Papailiopoulos]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2307.03381.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2307.03381"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

40. `https://arxiv.org/abs/2306.14308#google`:
    [`“Let’s Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning”`{=html}](https://arxiv.org/abs/2306.14308#google "'Let’s Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning', Ma et al 2023"){#ma-et-al-2023-2
    .link-annotated},
    [Xiao Ma, [Swaroop Mishra](https://swarooprm.github.io/ "Swaroop Mishra"){.link-annotated-partial
    .id-not}, Ahmad Beirami[, Alex Beutel, Jilin Chen]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2306.14308%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2306.14308#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

41. `https://arxiv.org/abs/2306.00323`:
    [`“Thought Cloning: Learning to Think While Acting by Imitating Human Thinking”`{=html}](https://arxiv.org/abs/2306.00323 "'Thought Cloning: Learning to Think while Acting by Imitating Human Thinking', Hu & Clune 2023"){#hu-clune-2023
    .link-annotated},
    [Shengran Hu, [Jeff Clune](http://jeffclune.com/ "Jeff Clune—Professor—Computer Science—University of British Columbia"){.link-annotated-partial
    .id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2306.00323.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2306.00323"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

42. `https://arxiv.org/abs/2305.20050#openai`:
    [`“Let’s Verify Step by Step”`{=html}](https://arxiv.org/abs/2305.20050#openai "'Let’s Verify Step by Step', Lightman et al 2023"){#lightman-et-al-2023
    .link-annotated},
    [Hunter Lightman, Vineet Kosaraju, Yura Burda[, Harri Edwards, Bowen Baker, Teddy Lee, [Jan Leike](https://jan.leike.name/ "Jan Leike"){.link-annotated-partial
    .id-not}, [John Schulman](http://joschu.net/ "'John Schulman’s Homepage', Schulman 2026"){.link-annotated-partial
    .id-not}, [Ilya Sutskever](https://en.wikipedia.org/wiki/Ilya_Sutskever "Ilya Sutskever"){.link-annotated-partial
    .id-not}, [Karl Cobbe](https://scholar.google.com/citations?user=stCljMYAAAAJ "Karl Cobbe"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2305.20050%2523openai.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2305.20050#openai"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

43. `https://arxiv.org/abs/2305.13534`:
    [`“How Language Model Hallucinations Can Snowball”`{=html}](https://arxiv.org/abs/2305.13534 "'How Language Model Hallucinations Can Snowball', Zhang et al 2023"){#zhang-et-al-2023-14
    .link-annotated},
    [Muru Zhang, Ofir Press, William Merrill[, Alisa Liu, [Noah Smith](https://en.wikipedia.org/wiki/Noah_Smith_(writer) "Noah Smith (writer)"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2305.13534.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2305.13534"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

44. `https://arxiv.org/abs/2305.10601#deepmind`:
    [`“Tree of Thoughts (ToT): Deliberate Problem Solving With Large Language Models”`{=html}](https://arxiv.org/abs/2305.10601#deepmind "'Tree of Thoughts (ToT): Deliberate Problem Solving with Large Language Models', Yao et al 2023"){#yao-et-al-2023
    .link-annotated},
    [Shunyu Yao, Dian Yu, Jeffrey Zhao[, Izhak Shafran, [Thomas L. Griffiths](https://en.wikipedia.org/wiki/Thomas_L._Griffiths "Thomas L. Griffiths"){.id-not}, [Yuan Cao](https://en.wikipedia.org/wiki/Yuan_Cao "Yuan Cao"){.id-not}, [Karthik Rajagopal Narasimhan](https://karthikncode.github.io/ "'Karthik R. Narasimhan homepage', Narasimhan 2026"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2305.10601%2523deepmind.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2305.10601#deepmind"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

45. `https://arxiv.org/abs/2305.04388`:
    [`“Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-Of-Thought Prompting”`{=html}](https://arxiv.org/abs/2305.04388 "'Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting', Turpin et al 2023"){#turpin-et-al-2023
    .link-annotated},
    [Miles Turpin, [Julian Michael](https://julianmichael.org/ "Julian Michael"){.link-annotated-partial
    .id-not}, [Ethan Perez](https://ethanperez.net/ "Ethan Perez"){.link-annotated-partial
    .id-not}, [Samuel R. Bowman](https://cims.nyu.edu/~sbowman/ "Sam Bowman"){.link-annotated-partial
    .id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2305.04388.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2305.04388"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

46. `https://arxiv.org/abs/2305.02301#google`:
    [`“Distilling Step-By-Step! Outperforming Larger Language Models With Less Training Data and Smaller Model Sizes”`{=html}](https://arxiv.org/abs/2305.02301#google "'Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes', Hsieh et al 2023"){#hsieh-et-al-2023-2
    .link-annotated},
    [Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh[, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, Tomas Pfister]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2305.02301%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2305.02301#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

47. `https://arxiv.org/abs/2304.11490`:
    [`“Boosting Theory-Of-Mind Performance in Large Language Models via Prompting”`{=html}](https://arxiv.org/abs/2304.11490 "'Boosting Theory-of-Mind Performance in Large Language Models via Prompting', Moghaddam & Honey 2023"){#moghaddam-honey-2023
    .link-annotated},
    [Shima Rahimi Moghaddam, Christopher J. Honey]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2304.11490.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2304.11490"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

48. `https://arxiv.org/abs/2304.02015#alibaba`:
    [`“How Well Do Large Language Models Perform in Arithmetic Tasks?”`{=html}](https://arxiv.org/abs/2304.02015#alibaba "'How well do Large Language Models perform in Arithmetic tasks?', Yuan et al 2023"){#yuan-et-al-2023-2
    .link-annotated},
    [Zheng Yuan, Hongyi Yuan, Chuanqi Tan[, [Wei Wang](https://web.cs.ucla.edu/~weiwang/ "Wei Wang’s Home Page"){.link-annotated-partial
    .id-not}, Songfang Huang]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2304.02015%2523alibaba.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2304.02015#alibaba"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

49. `https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335905`:
    [`“ChatGPT Goes to Law School”`{=html}](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335905 "'ChatGPT Goes to Law School', Choi et al 2023"){#choi-et-al-2023
    .link-annotated},
    [Jonathan H. Choi, Kristin E. Hickman, Amy Monahan, Daniel Schwarcz]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fpapers.ssrn.com%252Fsol3%252Fpapers.cfm%253Fabstract_id%253D4335905.html "Directory-tag link-bibliography for link https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335905"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

50. `https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335945`:
    [`“Large Language Models As Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards”`{=html}](https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335945 "'Large Language Models as Fiduciaries: A Case Study Toward Robustly Communicating With Artificial Intelligence Through Legal Standards', Nay 2023"){#nay-2023
    .link-annotated}, [John Nay]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fpapers.ssrn.com%252Fsol3%252Fpapers.cfm%253Fabstract_id%253D4335945.html "Directory-tag link-bibliography for link https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4335945"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

51. `https://arxiv.org/abs/2301.01751#elicit`:
    [`“Iterated Decomposition: Improving Science Q&amp;A by Supervising Reasoning Processes”`{=html}](https://arxiv.org/abs/2301.01751#elicit "'Iterated Decomposition: Improving Science Q&amp;A by Supervising Reasoning Processes', Reppert et al 2023"){#reppert-et-al-2023
    .link-annotated},
    [Justin Reppert, Ben Rachbach, Charlie George[, Luke Stebbing, Jungwon Byun, Maggie Appleton, Andreas Stuhlmüller]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2301.01751%2523elicit.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2301.01751#elicit"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

52. `https://arxiv.org/abs/2210.11399#google`:
    [`“U-PaLM: Transcending Scaling Laws With 0.1% Extra Compute”`{=html}](https://arxiv.org/abs/2210.11399#google "'U-PaLM: Transcending Scaling Laws with 0.1% Extra Compute', Tay et al 2022"){#tay-et-al-2022-upalm
    .link-annotated},
    [[Yi Tay](https://www.yitay.net/ "Yi Tay"){.link-annotated-partial
    .id-not}, [Jason Wei](https://www.jasonwei.net/ "Jason Wei"){.id-not}, Hyung Won Chung[, Vinh Q. Tran, David R. So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, [Denny Zhou](https://dennyzhou.github.io/ "Denny Zhou’s Home Page"){.link-annotated-partial
    .id-not}, [Donald Metzler](https://www.don-metzler.net/ "About—Donald Metzler"){.link-annotated-partial
    .id-not}, Slav Petrov, [Neil Houlsby](https://neilhoulsby.github.io/ "Neil Houlsby"){.link-annotated-partial
    .id-not}, [Quoc V. Le](https://en.wikipedia.org/wiki/Quoc_V._Le "Quoc V. Le"){.id-not}, [Mostafa Dehghani](https://scholar.google.com/citations?user=MiHOX3QAAAAJ "Mostafa Dehghani"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2210.11399%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2210.11399#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

53. `https://arxiv.org/abs/2210.11610#google`:
    [`“Large Language Models Can Self-Improve”`{=html}](https://arxiv.org/abs/2210.11610#google "'Large Language Models Can Self-Improve', Huang et al 2022"){#huang-et-al-2022-2
    .link-annotated},
    [Jiaxin Huang, [Shixiang Shane Gu](https://sites.google.com/view/gugurus/home){.id-not}, Le Hou[, Yuexin Wu, Xuezhi Wang, Hongkun Yu, [Jiawei Han](https://en.wikipedia.org/wiki/Jiawei_Han "Jiawei Han"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2210.11610%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2210.11610#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

54. `https://arxiv.org/abs/2210.09261#google`:
    [`“Challenging BIG-Bench Tasks (BBH) and Whether Chain-Of-Thought Can Solve Them”`{=html}](https://arxiv.org/abs/2210.09261#google "'Challenging BIG-Bench Tasks (BBH) and Whether Chain-of-Thought Can Solve Them', Suzgun et al 2022"){#suzgun-et-al-2022-1
    .link-annotated},
    [Mirac Suzgun, Nathan Scales, Nathanael Schärli[, Sebastian Gehrmann, [Yi Tay](https://www.yitay.net/ "Yi Tay"){.link-annotated-partial
    .id-not}, Hyung Won Chung, Aakanksha Chowdhery, [Quoc V. Le](https://en.wikipedia.org/wiki/Quoc_V._Le "Quoc V. Le"){.id-not}, [Ed H. Chi](https://en.wikipedia.org/wiki/Ed_H._Chi "Ed H. Chi"){.id-not}, [Denny Zhou](https://dennyzhou.github.io/ "Denny Zhou’s Home Page"){.link-annotated-partial
    .id-not}, [Jason Wei](https://www.jasonwei.net/ "Jason Wei"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2210.09261%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2210.09261#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

55. `https://arxiv.org/abs/2210.03350#allen`:
    [`“Self-Ask: Measuring and Narrowing the Compositionality Gap in Language Models (Bamboogle)”`{=html}](https://arxiv.org/abs/2210.03350#allen "'Self-Ask: Measuring and Narrowing the Compositionality Gap in Language Models (Bamboogle)', Press et al 2022"){#press-et-al-2022
    .link-annotated},
    [Ofir Press, Muru Zhang, Sewon Min[, [Ludwig Schmidt](https://en.wikipedia.org/wiki/Ludwig_Schmidt "Ludwig Schmidt"){.id-not}, [Noah A. Smith](https://nasmith.github.io/ "Noah A. Smith"){.link-annotated-partial
    .id-not}, [Mike Lewis](https://scholar.google.com/citations?user=SnQnQicAAAAJ "Mike Lewis"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2210.03350%2523allen.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2210.03350#allen"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

56. `https://arxiv.org/abs/2210.03057#google`:
    [`“Language Models Are Multilingual Chain-Of-Thought Reasoners”`{=html}](https://arxiv.org/abs/2210.03057#google "'Language Models are Multilingual Chain-of-Thought Reasoners', Shi et al 2022"){#shi-et-al-2022-2
    .link-annotated},
    [Freda Shi, Mirac Suzgun, Markus Freitag[, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, [Yi Tay](https://www.yitay.net/ "Yi Tay"){.link-annotated-partial
    .id-not}, Sebastian Ruder, [Denny Zhou](https://dennyzhou.github.io/ "Denny Zhou’s Home Page"){.link-annotated-partial
    .id-not}, Dipanjan Das, [Jason Wei](https://www.jasonwei.net/ "Jason Wei"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2210.03057%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2210.03057#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

57. `https://arxiv.org/abs/2210.03629#google`:
    [`“ReAct: Synergizing Reasoning and Acting in Language Models”`{=html}](https://arxiv.org/abs/2210.03629#google "'ReAct: Synergizing Reasoning and Acting in Language Models', Yao et al 2022"){#yao-et-al-2022-1
    .link-annotated},
    [Shunyu Yao, Jeffrey Zhao, Dian Yu[, Nan Du, Izhak Shafran, [Karthik Rajagopal Narasimhan](https://karthikncode.github.io/ "'Karthik R. Narasimhan homepage', Narasimhan 2026"){.id-not}, [Yuan Cao](https://en.wikipedia.org/wiki/Yuan_Cao "Yuan Cao"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2210.03629%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2210.03629#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

58. `https://arxiv.org/abs/2209.00840`:
    [`“FOLIO: Natural Language Reasoning With First-Order Logic”`{=html}](https://arxiv.org/abs/2209.00840 "'FOLIO: Natural Language Reasoning with First-Order Logic', Han et al 2022"){#han-et-al-2022
    .link-annotated},
    [Simeng Han, Hailey Schoelkopf, Yilun Zhao[, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, David Peng, Jonathan Fan, Yixin Liu, [Brian Wong](https://en.wikipedia.org/wiki/Brian_Wong "Brian Wong"){.id-not}, Malcolm Sailor, Ansong Ni, Linyong Nan, [Jungo Kasai](https://jungokasai.github.io/ "About me—Jungo Kasai"){.link-annotated-partial
    .id-not}, Tao Yu, Rui Zhang, Shafiq Joty, Alexander R. Fabbri, Wojciech Kryscinski, Xi Victoria Lin, [Caiming Xiong](http://cmxiong.com/ "Caiming Xiong—Home Page"){.link-annotated-partial
    .id-not}, [Dragomir Radev](https://en.wikipedia.org/wiki/Dragomir_Radev "Dragomir Radev"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2209.00840.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2209.00840"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

59. `https://arxiv.org/abs/2207.08143`:
    [`“Can Large Language Models Reason about Medical Questions?”`{=html}](https://arxiv.org/abs/2207.08143 "'Can large language models reason about medical questions?', Liévin et al 2022"){#liévin-et-al-2022
    .link-annotated},
    [Valentin Liévin, Christoffer Egeberg Hother, Ole Winther]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2207.08143.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2207.08143"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

60. `https://arxiv.org/abs/2207.05608#google`:
    [`“Inner Monologue: Embodied Reasoning through Planning With Language Models”`{=html}](https://arxiv.org/abs/2207.05608#google "'Inner Monologue: Embodied Reasoning through Planning with Language Models', Huang et al 2022"){#huang-et-al-2022-5
    .link-annotated},
    [Wenlong Huang, [Fei Xia](https://en.wikipedia.org/wiki/Fei_Xia "Fei Xia"){.id-not}, Ted Xiao[, Harris Chan, Jacky Liang, Pete Florence, [Andy Zeng](https://en.wikipedia.org/wiki/Andy_Zeng "Andy Zeng"){.id-not}, Jonathan Tompson, [Igor Mordatch](https://scholar.google.com/citations?user=Vzr1RukAAAAJ "Igor Mordatch"){.id-not}, Yevgen Chebotar, Pierre Sermanet, Noah Brown, Tomas Jackson, Linda Luu, [Sergey Levine](https://scholar.google.com/citations?user=8R35rCwAAAAJ "Sergey Levine"){.id-not}, [Karol Hausman](https://en.wikipedia.org/wiki/Karol_Hausman "Karol Hausman"){.id-not}, Brian Ichter]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2207.05608%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2207.05608#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

61. `https://arxiv.org/abs/2207.05221#anthropic`:
    [`“Language Models (Mostly) Know What They Know”`{=html}](https://arxiv.org/abs/2207.05221#anthropic "'Language Models (Mostly) Know What They Know', Kadavath et al 2022"){#kadavath-et-al-2022
    .link-annotated},
    [[Saurav Kadavath](https://scholar.google.com/citations?user=Z2Uo_FcAAAAJ "Saurav Kadavath"){.id-not}, Tom Conerly, [Amanda Askell](https://askell.io/ "About Me"){.link-annotated-partial
    .id-not}[, [Tom Henighan](https://www.tomhenighan.com/ "Tom Henighan"){.link-annotated-partial
    .id-not}, Dawn Drain, [Ethan Perez](https://ethanperez.net/ "Ethan Perez"){.link-annotated-partial
    .id-not}, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, [Scott Johnston](https://en.wikipedia.org/wiki/Scott_Johnston "Scott Johnston"){.id-not}, Sheer El-Showk, [Andy L. Jones](https://andyljones.com/ "andy jones"){.id-not}, [Nelson Elhage](https://nelhage.com/){.link-annotated-partial
    .id-not}, Tristan Hume, Anna Chen, [Yuntao Bai](https://scholar.google.com/citations?user=r7GUEVsAAAAJ "Yuntao Bai"){.id-not}, [Samuel R. Bowman](https://cims.nyu.edu/~sbowman/ "Sam Bowman"){.link-annotated-partial
    .id-not}, Stanislav Fort, [Deep Ganguli](https://dganguli.github.io/pweb/ "Deep Ganguli—Research Scientist at Anthropic"){.link-annotated-partial
    .id-not}, [Danny Hernandez](https://en.wikipedia.org/wiki/Danny_Hernandez "Danny Hernandez"){.id-not}, Josh Jacobson, [Jackson Kernion](https://jacksonkernion.com/){.link-annotated-partial
    .id-not}, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, [Dario Amodei](https://en.wikipedia.org/wiki/Dario_Amodei "Dario Amodei"){.link-annotated-partial
    .id-not}, [Tom B. Brown](https://scholar.google.com/citations?user=RLvsC94AAAAJ "Tom B. Brown"){.id-not}, [Jack Clark](https://jack-clark.net/about/){.link-annotated-partial
    .id-not}, [Nicholas Joseph](https://scholar.google.com/citations?user=Z5NeEv4AAAAJ "Nicholas Joseph"){.id-not}, [Ben Mann](https://scholar.google.com/citations?user=McBoXK0AAAAJ "Benjamin Mann"){.id-not}, [Sam McCandlish](https://scholar.google.com/citations?user=gHp0pu4AAAAJ "Sam McCandlish"){.id-not}, Chris Olah, [Jared Kaplan](https://sites.krieger.jhu.edu/jared-kaplan/ "Jared Kaplan"){.link-annotated-partial
    .id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2207.05221%2523anthropic.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2207.05221#anthropic"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

62. `https://arxiv.org/abs/2205.10625#google`:
    [`“Least-To-Most Prompting Enables Complex Reasoning in Large Language Models”`{=html}](https://arxiv.org/abs/2205.10625#google "'Least-to-Most Prompting Enables Complex Reasoning in Large Language Models', Zhou et al 2022"){#zhou-et-al-2022-1
    .link-annotated},
    [[Denny Zhou](https://dennyzhou.github.io/ "Denny Zhou’s Home Page"){.link-annotated-partial
    .id-not}, Nathanael Schärli, Le Hou[, [Jason Wei](https://www.jasonwei.net/ "Jason Wei"){.id-not}, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Olivier Bousquet, [Quoc V. Le](https://en.wikipedia.org/wiki/Quoc_V._Le "Quoc V. Le"){.id-not}, [Ed Chi](https://en.wikipedia.org/wiki/Ed_Chi "Ed Chi"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2205.10625%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2205.10625#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

63. `https://arxiv.org/abs/2205.09073#google`:
    [`“Dialog Inpainting: Turning Documents into Dialogues”`{=html}](https://arxiv.org/abs/2205.09073#google "'Dialog Inpainting: Turning Documents into Dialogues', Dai et al 2022"){#dai-et-al-2022-2
    .link-annotated},
    [Zhuyun Dai, Arun Tejasvi Chaganty, Vincent Zhao[, Aida Amini, Qazi Mamunur Rashid, Mike Green, Kelvin Guu]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2205.09073%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2205.09073#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

64. `https://arxiv.org/abs/2205.05131#google`:
    [`“UL2: Unifying Language Learning Paradigms”`{=html}](https://arxiv.org/abs/2205.05131#google "'UL2: Unifying Language Learning Paradigms', Tay et al 2022"){#tay-et-al-2022-ul2
    .link-annotated},
    [[Yi Tay](https://www.yitay.net/ "Yi Tay"){.link-annotated-partial
    .id-not}, [Mostafa Dehghani](https://scholar.google.com/citations?user=MiHOX3QAAAAJ "Mostafa Dehghani"){.id-not}, Vinh Q. Tran[, Xavier Garcia, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, [Neil Houlsby](https://neilhoulsby.github.io/ "Neil Houlsby"){.link-annotated-partial
    .id-not}, [Donald Metzler](https://www.don-metzler.net/ "About—Donald Metzler"){.link-annotated-partial
    .id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2205.05131%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2205.05131#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

65. `https://arxiv.org/abs/2204.00598#google`:
    [`“Socratic Models: Composing Zero-Shot Multimodal Reasoning With Language”`{=html}](https://arxiv.org/abs/2204.00598#google "'Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language', Zeng et al 2022"){#zeng-et-al-2022-2
    .link-annotated},
    [[Andy Zeng](https://en.wikipedia.org/wiki/Andy_Zeng "Andy Zeng"){.id-not}, Adrian Wong, Stefan Welker[, Krzysztof Choromanski, Federico Tombari, Aveek Purohit, Michael Ryoo, Vikas Sindhwani, Johnny Lee, Vincent Vanhoucke, Pete Florence]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2204.00598%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2204.00598#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

66. `https://arxiv.org/abs/2203.11171#google`:
    [`“Self-Consistency Improves Chain-Of-Thought Reasoning in Language Models”`{=html}](https://arxiv.org/abs/2203.11171#google "'Self-Consistency Improves Chain-of-Thought Reasoning in Language Models', Wang et al 2022"){#wang-et-al-2022-20
    .link-annotated},
    [Xuezhi Wang, [Jason Wei](https://www.jasonwei.net/ "Jason Wei"){.id-not}, Dale Schuurmans[, [Quoc V. Le](https://en.wikipedia.org/wiki/Quoc_V._Le "Quoc V. Le"){.id-not}, [Ed Chi](https://en.wikipedia.org/wiki/Ed_Chi "Ed Chi"){.id-not}, [Denny Zhou](https://dennyzhou.github.io/ "Denny Zhou’s Home Page"){.link-annotated-partial
    .id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2203.11171%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2203.11171#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

67. `https://arxiv.org/abs/2201.11903#google`:
    [`“Chain-Of-Thought Prompting Elicits Reasoning in Large Language Models”`{=html}](https://arxiv.org/abs/2201.11903#google "'Chain-of-Thought Prompting Elicits Reasoning in Large Language Models', Wei et al 2022"){#wei-et-al-2022-4
    .link-annotated},
    [[Jason Wei](https://www.jasonwei.net/ "Jason Wei"){.id-not}, Xuezhi Wang, Dale Schuurmans[, [Maarten Bosma](https://ma2rten.github.io/ "Maarten Bosma"){.link-annotated-partial
    .id-not}, [Ed Chi](https://en.wikipedia.org/wiki/Ed_Chi "Ed Chi"){.id-not}, [Quoc V. Le](https://en.wikipedia.org/wiki/Quoc_V._Le "Quoc V. Le"){.id-not}, [Denny Zhou](https://dennyzhou.github.io/ "Denny Zhou’s Home Page"){.link-annotated-partial
    .id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2201.11903%2523google.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2201.11903#google"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

68. `https://arxiv.org/abs/2201.11473#microsoft`:
    [`“Reasoning Like Program Executors”`{=html}](https://arxiv.org/abs/2201.11473#microsoft "'Reasoning Like Program Executors', Pi et al 2022"){#pi-et-al-2022
    .link-annotated},
    [Xinyu Pi, [Qian Liu](https://en.wikipedia.org/wiki/Qian_Liu "Qian Liu"){.id-not}, Bei Chen[, Morteza Ziyadi, Zeqi Lin, Yan Gao, Qiang Fu, Jian-Guang Lou, [Weizhu Chen](https://scholar.google.com/citations?user=LG_E-4EAAAAJ "Weizhu Chen"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2201.11473%2523microsoft.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2201.11473#microsoft"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

69. `https://arxiv.org/abs/2112.15594`:
    [`“A Neural Network Solves and Generates Mathematics Problems by Program Synthesis: Calculus, Differential Equations, Linear Algebra, and More”`{=html}](https://arxiv.org/abs/2112.15594 "'A Neural Network Solves and Generates Mathematics Problems by Program Synthesis: Calculus, Differential Equations, Linear Algebra, and More', Drori et al 2021"){#drori-et-al-2021
    .link-annotated},
    [Iddo Drori, Sunny Tran, Roman Wang[, Newman Cheng, Kevin Liu, Leonard Tang, Elizabeth Ke, Nikhil Singh, Taylor L. Patti, Jayson Lynch, Avi Shporer, [Nakul Verma](https://en.wikipedia.org/wiki/Nakul_Verma "Nakul Verma"){.id-not}, [Eugene Wu](https://en.wikipedia.org/wiki/Eugene_Wu "Eugene Wu"){.id-not}, [Gilbert Strang](https://en.wikipedia.org/wiki/Gilbert_Strang "Gilbert Strang"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2112.15594.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2112.15594"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

70. `https://openai.com/research/webgpt`:
    [`“WebGPT: Improving the Factual Accuracy of Language Models through Web Browsing”`{=html}](https://openai.com/research/webgpt "'WebGPT: Improving the factual accuracy of language models through web browsing', Hilton et al 2021"){#hilton-et-al-2021-1
    .link-annotated},
    [[Jacob Hilton](https://www.jacobh.co.uk/ "Jacob Hilton's Homepage"){.link-annotated-partial
    .id-not}, Suchir Balaji, Reiichiro Nakano, [John Schulman](http://joschu.net/ "'John Schulman’s Homepage', Schulman 2026"){.link-annotated-partial
    .id-not}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fopenai.com%252Fresearch%252Fwebgpt.html "Directory-tag link-bibliography for link https://openai.com/research/webgpt"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

71. `https://arxiv.org/abs/2110.14168#openai`:
    [`“Training Verifiers to Solve Math Word Problems”`{=html}](https://arxiv.org/abs/2110.14168#openai "'Training Verifiers to Solve Math Word Problems', Cobbe et al 2021"){#cobbe-et-al-2021
    .link-annotated},
    [[Karl Cobbe](https://scholar.google.com/citations?user=stCljMYAAAAJ "Karl Cobbe"){.id-not}, Vineet Kosaraju, Mohammad Bavarian[, [Jacob Hilton](https://www.jacobh.co.uk/ "Jacob Hilton's Homepage"){.link-annotated-partial
    .id-not}, Reiichiro Nakano, [Christopher Hesse](https://scholar.google.com/citations?user=SgbbTp4AAAAJ "Christopher Hesse"){.id-not}, [John Schulman](http://joschu.net/ "'John Schulman’s Homepage', Schulman 2026"){.link-annotated-partial
    .id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Farxiv.org%252Fabs%252F2110.14168%2523openai.html "Directory-tag link-bibliography for link https://arxiv.org/abs/2110.14168#openai"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

72. `https://sites.google.com/berkeley.edu/decision-transformer`:
    [`“Decision Transformer: Reinforcement Learning via Sequence Modeling”`{=html}](https://sites.google.com/berkeley.edu/decision-transformer "'Decision Transformer: Reinforcement Learning via Sequence Modeling', Chen et al 2021"){#decisiontransformer-blog
    .link-annotated},
    [[Lili Chen](https://www.lilichen.me/ "Lili Chen"){.link-annotated-partial
    .id-not}, [Kevin Lu](https://kevinlu.ai/ "Kevin Lu"){.link-annotated-partial
    .id-not}, [Aravind Rajeswaran](https://aravindr93.github.io/ "Aravind Rajeswaran"){.link-annotated-partial
    .id-not}[, [Kimin Lee](https://sites.google.com/view/kiminlee "Kimin Lee"){.id-not}, [Aditya Grover](https://scholar.google.com/citations?user=oOhnPUgAAAAJ "Aditya Grover"){.id-not}, [Michael Laskin](https://scholar.google.com/citations?user=DOGDnwsAAAAJ "Michael (misha) Laskin"){.id-not}, [Pieter Abbeel](https://en.wikipedia.org/wiki/Pieter_Abbeel "Pieter Abbeel"){.id-not}, [Aravind Srinivas](https://scholar.google.com/citations?user=GhrKC1gAAAAJ "Aravind Srinivas"){.id-not}, [Igor Mordatch](https://scholar.google.com/citations?user=Vzr1RukAAAAJ "Igor Mordatch"){.id-not}]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fsites.google.com%252Fberkeley.edu%252Fdecision-transformer.html "Directory-tag link-bibliography for link https://sites.google.com/berkeley.edu/decision-transformer"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

73. `https://gptprompts.wikidot.com/linguistics:word-in-context#toc3`:
    [`“Word in Context: Agent and Agent Clarification (69% Dev)”`{=html}](https://gptprompts.wikidot.com/linguistics:word-in-context#toc3 "'Word in Context: Agent and Agent Clarification (69% Dev)', Brockman 2020"){#brockman-2020
    .link-annotated}, [Matt Brockman]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fgptprompts.wikidot.com%252Flinguistics%253Aword-in-context%2523toc3.html "Directory-tag link-bibliography for link https://gptprompts.wikidot.com/linguistics:word-in-context#toc3"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

74. `https://news.ycombinator.com/item?id=23990902`:
    [`“I Found That Getting GPT-3 to Add Its Own "Internal Monologue" in Parentheses to Be a Helpful Strategy…”`{=html}](https://news.ycombinator.com/item?id=23990902 "'I found that getting GPT-3 to add its own "internal monologue" in parentheses to be a helpful strategy…', blixt 2020"){#blixt-2020
    .link-annotated}, [blixt]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fnews.ycombinator.com%252Fitem%253Fid%253D23990902.html "Directory-tag link-bibliography for link https://news.ycombinator.com/item?id=23990902"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

75. `https://x.com/kleptid/status/1284069270603866113`:
    [`“Seems to Work”`{=html}](https://x.com/kleptid/status/1284069270603866113 "'Seems to work', KaryoKleptid 2020"){#karyokleptid-2020-2
    .link-annotated}, [KaryoKleptid]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fx.com%252Fkleptid%252Fstatus%252F1284069270603866113.html "Directory-tag link-bibliography for link https://x.com/kleptid/status/1284069270603866113"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

76. `https://x.com/kleptid/status/1284098635689611264`:
    [`“Teaching GPT-3 to Do a Brute Force ‘For Loop’ Checking Answers Also Seems to Work”`{=html}](https://x.com/kleptid/status/1284098635689611264 "'Teaching GPT-3 to do a brute force ‘for loop’ checking answers also seems to work', KaryoKleptid 2020"){#karyokleptid-2020-1
    .link-annotated}, [KaryoKleptid]{.author .cite-author}

    [[link-bibliography](/metadata/annotation/link-bibliography/https%253A%252F%252Fx.com%252Fkleptid%252Fstatus%252F1284098635689611264.html "Directory-tag link-bibliography for link https://x.com/kleptid/status/1284098635689611264"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}

77. `2018-bisra.pdf`:
    [`“Inducing Self-Explanation: a Meta-Analysis”`{=html}](/doc/psychology/spaced-repetition/2018-bisra.pdf "'Inducing Self-Explanation: a Meta-Analysis', Bisra et al 2018"){#bisra-et-al-2018
    .link-annotated},
    [Kiran Bisra, Qing Liu, John C. Nesbit[, Farimah Salimi, Philip H. Winne]{.collapse}]{.author}

    [[link-bibliography](/metadata/annotation/link-bibliography/%252Fdoc%252Fpsychology%252Fspaced-repetition%252F2018-bisra.pdf.html "Directory-tag link-bibliography for link /doc/psychology/spaced-repetition/2018-bisra.pdf"){.id-not
    .include-even-when-collapsed}]{.collapse
    .tag-index-link-bibliography-block}
