Legacy Atlas

Mathematical Foundations

The foundations library is the broad map: 100 interconnected concepts, legacy demos, canonical papers, and migration links into the newer domain notebooks.

100 concepts37 demos336 connections13 learning phases
1Core
2Optim
3Gen
4Rep
5Scale
6Systems
lab surfacepredict, drag, test

How To Use This Layer

Foundations is the wide map; domains are the refined notebooks.

Keep the old map useful while migration continues: browse broadly here, then step into domain notebooks when a concept has a newer treatment.
SurveyUse the graph when you need the map.

The legacy foundations graph is still the fastest way to see prerequisites, hubs, and semantic edges across the full 100-concept atlas.

Open concept map
SequenceUse the study path when order matters.

The phases turn a broad topic list into a teachable route from core training ideas to systems, alignment, and frontier methods.

Follow phases
LabUse demo filters when you want to learn by doing.

The old dark demo surfaces stay intact here while the surrounding experience points learners toward notebooks and domain pages.

Find demos

Recommended Study Order

Build understanding from fundamentals to frontier techniques. Each phase builds on the previous one.

1

Core probabilistic training + transformers

All 100 Concepts

Search and filter to find what you're looking for.

1โ„’

ML/CE/KL

Maximum Likelihood, Cross-Entropy & KL Divergence

Core Trainingโญ 261 paperโ–ถ Demo
Explore โ†’
2โŠ—

Attention

Scaled Dot-Product Attention & Transformer Layers

Core Trainingโญ 231 paperโ–ถ Demo
Explore โ†’
3โˆ‡

Adam

Adam & Adaptive Gradient Methods

Optimization๐Ÿ”— 42 papersโ–ถ Demo
Explore โ†’
4โŒ‡

Sharpness

Loss Landscapes, Sharpness & Flat Minima

Optimizationโญ 112 papersโ–ถ Demo
Explore โ†’
5โˆช

Double Descent

Overparameterization & Generalization, Double Descent

Optimization๐Ÿ”— 42 papersโ–ถ Demo
Explore โ†’
6ฮ˜

NTK

Neural Tangent Kernel & Infinite-Width Limits

Theory๐Ÿ”— 41 paperโ–ถ Demo
Explore โ†’
7โ„ค

VAEs

Variational Autoencoders & Variational Inference

Generative Models๐Ÿ”— 51 paperโ–ถ Demo
Explore โ†’
8โš”

GANs

GANs & Adversarial Divergence Minimization

Generative Models๐Ÿ”— 22 papersโ–ถ Demo
Explore โ†’
9โˆ‚

Diffusion

Diffusion, Score-Based Models & Flow Matching

Generative Modelsโญ 114 papersโ–ถ Demo
Explore โ†’
10โ—Ž

Embeddings

Representation Learning & Embedding Geometry

Representationsโญ 141 paperโ–ถ Demo
Explore โ†’
11โŠ•

Superposition

Superposition, Sparse Features & Monosemanticity

Representations๐Ÿ”— 32 papersโ–ถ Demo
Explore โ†’
12โšฒ

Probing

Probing, Linear Classifier Probes & Activation Analysis

Representations๐Ÿ”— 52 papersโ–ถ Demo
Explore โ†’
13โŠ›

Circuits

Transformer Circuits, Induction Heads & Mechanistic Interpretability

Representations๐Ÿ”— 72 papersโ–ถ Demo
Explore โ†’
14โ†—

Scaling

Scaling Laws & Emergent Abilities

Scaling & Alignment๐Ÿ”— 73 papersโ–ถ Demo
Explore โ†’
15โš–

RLHF

Preference-Based Alignment: RLHF, Reward Modeling, Constitutional AI

Scaling & Alignmentโญ 173 papersโ–ถ Demo
Explore โ†’
16โšก

Efficiency

Efficiency: Quantization, Distillation, LoRA & Sparse MoE

Efficiency๐Ÿ”— 93 papersโ–ถ Demo
Explore โ†’
17โˆ€

Theory

Theoretical Foundations: PAC Learning, MDL & Information Bottleneck

Theory๐Ÿ”— 93 papersโ–ถ Demo
Explore โ†’
19โš™

Efficient Attention

Efficient Attention at Scale: KV Cache, GQA & FlashAttention

Efficiencyโญ 113 papersโ–ถ Demo
Explore โ†’
18โ†ป

RoPE

Rotary Position Embeddings (RoPE)

Representations๐Ÿ”— 33 papersโ–ถ Demo
Explore โ†’
20โฉ

Speculative Decoding

Speculative Decoding: Lossless Multi-Token Generation

Efficiency๐Ÿ”— 63 papersโ–ถ Demo
Explore โ†’
21โšก

LLM Serving

LLM Serving at Scale: Prefill, Decode & Continuous Batching

Efficiency๐Ÿ”— 73 papersโ–ถ Demo
Explore โ†’
22๐Ÿ”€

MoE

Sparse Mixture of Experts: Routing, Load Balancing & Expert Parallelism

Efficiency๐Ÿ”— 83 papersโ–ถ Demo
Explore โ†’
23โšก

MoE Serving

MoE Serving & Scheduling: Token Dispatch, All-to-All, Disaggregated Parallelism

Efficiency๐Ÿ”— 43 papersโ–ถ Demo
Explore โ†’
24๐ŸŽฏ

DPO

Direct Preference Optimization: RL-Free Alignment from Human Preferences

Scaling & Alignment๐Ÿ”— 53 papersโ–ถ Demo
Explore โ†’
25๐Ÿ‘

KTO

KTO: Alignment from Binary Feedback via Human-Aware Losses

Scaling & Alignment๐Ÿ”— 33 papersโ–ถ Demo
Explore โ†’
26โš ๏ธ

Reward Hacking

Reward Hacking & Overoptimization: Goodhart's Law in Preference Optimization

Scaling & Alignment๐Ÿ”— 33 papersโ–ถ Demo
Explore โ†’
27๐Ÿ”

Sparse Autoencoders

Sparse Autoencoders at Scale: Feature Dictionaries for Mechanistic Interpretability

Representations๐Ÿ”— 63 papersโ–ถ Demo
Explore โ†’
28๐Ÿ”ฌ

Circuit Discovery

Automated Circuit Discovery: Patching, Attribution & Decomposition at Scale

Representations๐Ÿ”— 43 papersโ–ถ Demo
Explore โ†’
29๐ŸŽš๏ธ

Activation Steering

Activation Steering: Feature-Guided Interventions for Inference-Time Control

Representations๐Ÿ”— 43 papersโ–ถ Demo
Explore โ†’
30๐Ÿ“

Long Context

Long Context Engineering: RoPE Scaling, KV Compression & Memory Optimization

Efficiencyโญ 103 papersโ–ถ Demo
Explore โ†’
31๐Ÿ”€

SSMs & Hybrids

State Space Models & Hybrid Architectures: Mamba-2, Jamba, Griffin

Core Training๐Ÿ”— 53 papersโ–ถ Demo
Explore โ†’
32๐Ÿ–ผ๏ธ

Multimodal VLP

Multimodal Foundations: Vision Encoders, Contrastive Learning & Cross-Attention Fusion

Representations๐Ÿ”— 53 papersโ–ถ Demo
Explore โ†’
33๐Ÿ”ค

Tokens

Tokenization & Vocabulary Design

Representations๐Ÿ”— 43 papersโ–ถ Demo
Explore โ†’
34๐ŸŽฒ

Decoding

Decoding & Sampling: Temperature, Top-p & Inference-Time Control

Core Training๐Ÿ”— 53 papersโ–ถ Demo
Explore โ†’
35โŸฒ

Backprop

Backpropagation & Automatic Differentiation

Optimization๐Ÿ”— 42 papers
Explore โ†’
36โˆ‡p

Score Matching

Score Matching & Score-Based Generative Models

Generative Models๐Ÿ”— 52 papers
Explore โ†’
37๐Ÿ“

ICL

In-Context Learning: Learning Without Weight Updates

Representations๐Ÿ”— 52 papers
Explore โ†’
38โš–๏ธ

OT/Wasserstein

Optimal Transport & Wasserstein Distance

Theory๐Ÿ”— 32 papers
Explore โ†’
39๐ŸŒŠ

Flows

Normalizing Flows: Exact Likelihood via Invertible Transforms

Generative Models๐Ÿ”— 32 papers
Explore โ†’
40๐ŸŽฏ

PPO

PPO: Proximal Policy Optimization

Scaling & Alignment๐Ÿ”— 11 paper
Explore โ†’
41โŠ•

Residuals

Residual Connections & Skip Connections

Core Training๐Ÿ”— 21 paper
Explore โ†’
42๐ŸŽš๏ธ

CFG

Classifier-Free Guidance in Diffusion

Generative Models๐Ÿ”— 21 paper
Explore โ†’
43๐Ÿ”

RAG

Retrieval-Augmented Generation (RAG)

Representations๐Ÿ”— 31 paper
Explore โ†’
44โš”๏ธ

Adversarial

Adversarial Examples & Robustness

Theory๐Ÿ”— 31 paper
Explore โ†’
45๐Ÿ’ก

Grokking

Grokking: Delayed Generalization

Theory๐Ÿ”— 31 paper
Explore โ†’
46๐Ÿ”ฌ

Logit Lens

Logit Lens: Probing Intermediate Representations

Representations๐Ÿ”— 41 paper
Explore โ†’
47๐Ÿ“‰

LR Schedules

Learning Rate Schedules: Warmup, Decay & Cycling

Optimization๐Ÿ”— 21 paper
Explore โ†’
48๐ŸŽฒ

Init

Weight Initialization: Xavier, He & ยตP

Optimization๐Ÿ”— 32 papers
Explore โ†’
49๐Ÿ”—

Contrastive

Contrastive Learning & InfoNCE

Representations๐Ÿ”— 22 papers
Explore โ†’
50๐ŸŒ

Distributed

Distributed Training: Data, Tensor & Pipeline Parallelism

Efficiency๐Ÿ”— 32 papers
Explore โ†’
51๐ŸŒณ

Beam Search

Beam Search & Structured Decoding

Core Training๐Ÿ”— 21 paper
Explore โ†’
52๐Ÿ’ง

Dropout

Dropout: Stochastic Regularization

Optimization๐Ÿ”— 21 paper
Explore โ†’
53โšก

EBMs

Energy-Based Models & Score Functions

Generative Models๐Ÿ”— 21 paper
Explore โ†’
54๐Ÿ“

LayerNorm

Layer Normalization & RMSNorm

Core Training๐Ÿ”— 22 papers
Explore โ†’
55โ„

Fisher Info

Fisher Information & Information Geometry

Theory๐Ÿ”— 42 papers
Explore โ†’
56๐Ÿงญ

Natural Grad

Natural Gradient & Riemannian Optimization

Optimization๐Ÿ”— 32 papers
Explore โ†’
57๐Ÿƒ

SGD+Momentum

SGD & Momentum: The Workhorses of Optimization

Optimization๐Ÿ”— 22 papers
Explore โ†’
58โš–๏ธ

AdamW

Weight Decay & AdamW: Decoupled Regularization

Optimization๐Ÿ”— 22 papers
Explore โ†’
59โœ‚๏ธ

Grad Clip

Gradient Clipping & Explosion Prevention

Optimization๐Ÿ”— 22 papers
Explore โ†’
60๐ŸŽฏ

Label Smooth

Label Smoothing & Soft Targets

Optimization๐Ÿ”— 22 papers
Explore โ†’
61๐Ÿ“Š

BatchNorm

Batch Normalization

Core Training๐Ÿ”— 12 papers
Explore โ†’
62๐Ÿงช

Distillation

Knowledge Distillation: Learning from Teachers

Efficiency๐Ÿ”— 62 papers
Explore โ†’
63๐Ÿ”ข

Quantization

Quantization: Compressing Models to Integers

Efficiency๐Ÿ”— 22 papers
Explore โ†’
64โœ‚๏ธ

Pruning

Pruning: Removing Unnecessary Weights

Efficiency๐Ÿ”— 22 papers
Explore โ†’
65๐Ÿ”„

SSL

Self-Supervised Learning: Labels from Structure

Representations๐Ÿ”— 32 papers
Explore โ†’
66๐ŸŒก๏ธ

Calibration

Calibration & Temperature Scaling

Theory๐Ÿ”— 32 papers
Explore โ†’
67๐Ÿšช

SwiGLU

SwiGLU & Gated Activations

Core Training๐Ÿ”— 12 papersโ–ถ Demo
Explore โ†’
68โšก

FlashAttn

FlashAttention: IO-Aware Attention

Efficiency๐Ÿ”— 22 papers
Explore โ†’
69๐Ÿ“œ

Constitutional

Constitutional AI: Principles-Based Alignment

Scaling & Alignment๐Ÿ”— 62 papers
Explore โ†’
70๐Ÿชž

Bregman

Bregman Divergence & Mirror Descent

Theory๐Ÿ”— 22 papers
Explore โ†’
71๐ŸŽญ

RKHS

Reproducing Kernel Hilbert Spaces

Theory๐Ÿ”— 22 papers
Explore โ†’
72๐Ÿ•ณ๏ธ

TDA

Persistent Homology & Topological Data Analysis

Theory๐Ÿ”— 22 papers
Explore โ†’
73๐Ÿ”€

Lie Groups

Lie Groups & Equivariant Networks

Theory๐Ÿ”— 22 papers
Explore โ†’
74๐Ÿง 

Test-Time

Test-Time Compute & Inference Scaling

Scaling & Alignment๐Ÿ”— 52 papers
Explore โ†’
75๐Ÿ’ญ

CoT

Chain-of-Thought Prompting

Scaling & Alignment๐Ÿ”— 42 papers
Explore โ†’
76๐ŸŒ

World Models

World Models & Model-Based RL

Theory๐Ÿ”— 22 papers
Explore โ†’
77๐Ÿญ

Synth Data

Synthetic Data & Self-Improvement

Scaling & Alignment๐Ÿ”— 22 papers
Explore โ†’
78๐ŸŽฏ

Consistency

Consistency Models: One-Step Diffusion

Generative Models๐Ÿ”— 22 papers
Explore โ†’
79๐Ÿ’พ

Checkpointing

Activation Checkpointing & Memory Efficiency

Efficiency๐Ÿ”— 22 papers
Explore โ†’
80๐Ÿ‘ฅ

GQA

Grouped Query Attention (GQA)

Efficiency๐Ÿ”— 21 paperโ–ถ Demo
Explore โ†’
81๐Ÿ“‹

PRMs

Process Reward Models

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
82๐Ÿค–

RLAIF

RLAIF: AI Feedback

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
83๐ŸŒŠ

Flow Match

Flow Matching & Rectified Flows

Generative Models๐Ÿ”— 21 paper
Explore โ†’
84๐Ÿ“

Instruct

Instruction Tuning

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
85๐ŸŽฏ

Deliberative

Deliberative Alignment

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
86โš”๏ธ

Debate

AI Safety via Debate

Scaling & Alignment๐Ÿ”— 11 paper
Explore โ†’
87๐Ÿ”„

IDA

Iterated Amplification

Scaling & Alignment๐Ÿ”— 31 paper
Explore โ†’
88๐Ÿ’ช

Weakโ†’Strong

Weak-to-Strong Generalization

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
89๐Ÿ”ด

Auto RedTeam

Automated Red Teaming

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
90๐Ÿ”ฎ

Mesa-Opt

Mesa-Optimization & Inner Alignment

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
91๐Ÿ˜ด

Sleepers

Sleeper Agents & Alignment Faking

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
92โš–๏ธ

LLM-as-Judge

Model-Graded Evaluations

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
93๐Ÿ”

Elicitation

Capability Elicitation & ELK

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
94๐Ÿฅช

Sandwich

Sandwiching Evaluations

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
95๐Ÿ“Š

MoD

Mixture-of-Depths

Efficiency๐Ÿ”— 21 paper
Explore โ†’
96๐ŸŒฒ

MCTS-LLM

Tree Search over Thoughts

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
97๐ŸŽฌ

VideoWM

Video World Models

Generative Models๐Ÿ”— 21 paper
Explore โ†’
98๐Ÿ”„

Self-Improve

Self-Improvement & Distillation Loops

Scaling & Alignment๐Ÿ”— 21 paper
Explore โ†’
99๐Ÿ“‰

Collapse

Model Collapse & Synthetic Data

Theory๐Ÿ”— 21 paper
Explore โ†’
100โ™พ๏ธ

InfCtx

Infinite Context Architectures

Efficiency๐Ÿ”— 21 paper
Explore โ†’

Why These 100 Concepts?

Complete Coverage

Together, these concepts explain the core mechanisms behind language models, diffusion models, and multimodal systems.

Missing Intuition

Each concept includes what's still poorly explained in textbooks and papers - the intuition gaps we aim to fill.

Connected Knowledge

See how concepts build on each other. Understand prerequisites and what each idea unlocks.