Shelf
Last updated: October 11, 2026
Papers and technical posts that I have enjoyed or thought about a lot:
- An Opinionated Guide to ML Research by John Schulman
- “A method that slightly improves on the baseline better be very simple, otherwise no one will bother using it — not even you.”
- The Big Blob of Compute Hypothesis by Dario Amodei
- Notably predates the pithier Bitter Lesson by Richard Sutton but lands on (a more thoughtfully caveated form of) the same intuition.
- Scaling Laws for Neural Language Models by Kaplan et al.
- Pretty ambitious to treat neural networks as subject to laws reminiscent of those applied to physical systems.
- Reader-friendly intro to LLM scaling laws: Three Kuhnian Revolutions in ML Training by Trevor Chow.
- A Neural Scaling Law from the Dimension of the Data Manifold by Sharma & Kaplan
- There’s something remarkable about the fact that almost every architectural improvement is a compute multiplier, and only the data distribution ever changes the exponent in log-log scaling plots.
- DeepSeek-R1 by Guo et al.
- Inadequate Equilibria by Eliezer Yudkowsky
- Reasoning from the efficient markets assumption (and knowing when it breaks) is important for deciding which research questions to prioritize.
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer by Yang et al.
- Pretty cool demonstration that a relatively mathematically sophisticated method can work shockingly well practically. Generally, galaxy-brained methods fail in deep learning.
- Coconut: Training Large Language Models to Reason in a Continuous Latent Space by Hao et al.
- Like the looped transformer which was vindicated on GPT-6 Astra, Coconut is a scaling axis that will eventually obviously work, IMO.
- Failures of Gradient-Based Deep Learning by Shalev-Shwartz et al.
- The failure of gradient-based neural networks to learn parities when the signal in the gradient is lower than the floating-point floor is suggestive when thinking about the limits of evolution in biological systems.
- Progress measures for grokking via mechanistic interpretability by Nanda et al.
- The most impressive reverse-engineering in mechanistic interpretability to date.
- Understanding a Law Firm through Study by Engram
- Synthetic midtraining data and RL environment generation feels destined to solve continual learning, and C&H is a great case study.
- GBA Eval by Stephen Yang
- Great tutorial and considerations on creating a fair evaluation (and RL environment) for frontier LLMs.
- You should be paranoid about what your eval is “actually” measuring. Other reading in that vein:
- About 30% of Humanity’s Last Exam Chemistry/biology Answers are Likely Wrong by Skarlinski et al.
- reward hacking in ether0 by Andrew White
- Is ProgramBench Impossible? by Saul Fuhrmann
- MirrorCode: AI can rebuild entire programs from behavior alone by Adamczewski et al., specifically Appendix E, “Comparison with ProgramBench”
- Computers can be understood by Nelson Elhage
- Quiet-STaR by Zelikman et al.
- As we near a data wall, this is a fruitful scaling axis to explore. I expect synthetic data generation to solve the bottleneck in data-limited regimes before Quiet-STaR-inspired methods do, though.
- Reinforcement Pre-Training by Dong et al. is the modern and possibly more practical version.
- Stealing Part of a Production Language Model by Carlini et al.
- Drives home that you can learn a lot about a served LLM from very little, alongside these papers:
- The Worst (But Only) Claude 3 Tokenizer by Javier Rando
- Stealing Reasoning Traces from Proprietary LLM APIs by Panfilov et al.
- Long-context latency scales quadratically for GPT-5.6 but nearly linearly for Claude 5 by Jason Li
- Drives home that you can learn a lot about a served LLM from very little, alongside these papers:
- Intriguing properties of neural networks by Szegedy et al.
- Muon: An optimizer for hidden layers in neural networks by Keller Jordan
- For an optimizer that finally beat Adam, Muon’s geometric intuition is remarkably simple.
- Generative design of bacteriophages with genome language models by King et al.
- MuZero: Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model by Schrittwieser et al.
- Reward is enough by Silver et al.
- SolidGoldMagikarp (plus, prompt generation) by Rumbelow & Watkins
- Watching a frontier LLM (Opus 5-tier) jerk out of assistant mode for yourself when fed some innocuous string is a jolt of an experience.
- The innocent gene by Joe Carlsmith
- Suggests the notion of the “innocent LLM” as a frame for thinking about misalignment.
- Specification gaming examples in AI by Victoria Krakovna
- the void by nostalgebraist
- Training AI to Paint with Code by Surya Nareddi
- Technical Report on the Pangram AI-Generated Text Classifier by Emi & Spero
- An AI-generated text detector with a false positive rate as low as Pangram’s is a remarkable achievement, and hard negative mining is a delightfully clever technique.
- How Our Data Shaped Neural Architecture Discovery, and How Automation Can Reshape the Future by Ehsan Amid
- Genie: Generative Interactive Environments by Bruce et al.
- Yan: Foundational Interactive Video Generation by Ye et al. is the modern version that (sanely) replaces the latent action model with ground-truth labels from synthetic environments.
- Catching crumbs from the table by Ted Chiang
- Refusal in Language Models Is Mediated by a Single Direction by Arditi et al.
- Neat demonstration of how shallow safety training can be. It gave the world “abliteration”, which is surprisingly effective.
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks is the other notable case study of how effective detection methods based on linear directions can be.
- Reading list (Arc.net version) put together for John Carmack by Ilya Sutskever
- The Art of Scaling Reinforcement Learning Compute for LLMs by Khatri et al.
- The conceptual decision to fit RL scaling as a sigmoid and the disciplined ablation of Every Tweak Possible are great contributions.
- Cheap RL tasks will waste compute by Erdil et al.
- Accurate structure prediction of biomolecular interactions with AlphaFold 3 by Abramson et al.
- Agents’ Last Exam by Sun et al.
- Compute Optimal Tokenization by Limisiewicz et al.
- Notable to me mostly for showing that heuristic-based byte-pair encoding is awkward and should eventually go.
- Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model by Ling Team
- Transformer Architecture: The Positional Encoding by Amirhossein Kazemnejad
- The sinusoidal position embedding is a really clever way to inject positional information into sequences without using parameters.
- ZeRO: Memory Optimizations Toward Training Trillion Parameter Models by Rajbhandari et al.
- ZeRO-2 is a free lunch.
- Efficient Memory Management for Large Language Model Serving with PagedAttention by Kwon et al. is another pleasingly simple LLM systems optimization.
- Interaction Models: A Scalable Approach to Human-AI Collaboration by Thinking Machines
- I reverse-engineered Instinct’s memory. Here’s exactly how it works by Dhravya Shah
- Instinct’s git- and filesystem-based memory system is ingenious.