The legacy foundations graph is still the fastest way to see prerequisites, hubs, and semantic edges across the full 100-concept atlas.
Open concept mapLegacy Atlas
Mathematical Foundations
The foundations library is the broad map: 100 interconnected concepts, legacy demos, canonical papers, and migration links into the newer domain notebooks.
How To Use This Layer
Foundations is the wide map; domains are the refined notebooks.
Keep the old map useful while migration continues: browse broadly here, then step into domain notebooks when a concept has a newer treatment.The phases turn a broad topic list into a teachable route from core training ideas to systems, alignment, and frontier methods.
Follow phasesThe old dark demo surfaces stay intact here while the surrounding experience points learners toward notebooks and domain pages.
Find demosRecommended Study Order
Build understanding from fundamentals to frontier techniques. Each phase builds on the previous one.
Optimization & generalization
Generative modeling families
Representation & interpretability
Modern efficiency & inference
Alignment & RLHF
Scaling, theory & multimodal
Advanced architectures & generation
Mathematical foundations & information geometry
Frontier research & scaling
Advanced alignment & safety research
All 100 Concepts
Search and filter to find what you're looking for.
ML/CE/KL
Maximum Likelihood, Cross-Entropy & KL Divergence
Explore โAttention
Scaled Dot-Product Attention & Transformer Layers
Explore โAdam
Adam & Adaptive Gradient Methods
Explore โSharpness
Loss Landscapes, Sharpness & Flat Minima
Explore โDouble Descent
Overparameterization & Generalization, Double Descent
Explore โNTK
Neural Tangent Kernel & Infinite-Width Limits
Explore โVAEs
Variational Autoencoders & Variational Inference
Explore โGANs
GANs & Adversarial Divergence Minimization
Explore โDiffusion
Diffusion, Score-Based Models & Flow Matching
Explore โEmbeddings
Representation Learning & Embedding Geometry
Explore โSuperposition
Superposition, Sparse Features & Monosemanticity
Explore โProbing
Probing, Linear Classifier Probes & Activation Analysis
Explore โCircuits
Transformer Circuits, Induction Heads & Mechanistic Interpretability
Explore โScaling
Scaling Laws & Emergent Abilities
Explore โRLHF
Preference-Based Alignment: RLHF, Reward Modeling, Constitutional AI
Explore โEfficiency
Efficiency: Quantization, Distillation, LoRA & Sparse MoE
Explore โTheory
Theoretical Foundations: PAC Learning, MDL & Information Bottleneck
Explore โEfficient Attention
Efficient Attention at Scale: KV Cache, GQA & FlashAttention
Explore โRoPE
Rotary Position Embeddings (RoPE)
Explore โSpeculative Decoding
Speculative Decoding: Lossless Multi-Token Generation
Explore โLLM Serving
LLM Serving at Scale: Prefill, Decode & Continuous Batching
Explore โMoE
Sparse Mixture of Experts: Routing, Load Balancing & Expert Parallelism
Explore โMoE Serving
MoE Serving & Scheduling: Token Dispatch, All-to-All, Disaggregated Parallelism
Explore โDPO
Direct Preference Optimization: RL-Free Alignment from Human Preferences
Explore โKTO
KTO: Alignment from Binary Feedback via Human-Aware Losses
Explore โReward Hacking
Reward Hacking & Overoptimization: Goodhart's Law in Preference Optimization
Explore โSparse Autoencoders
Sparse Autoencoders at Scale: Feature Dictionaries for Mechanistic Interpretability
Explore โCircuit Discovery
Automated Circuit Discovery: Patching, Attribution & Decomposition at Scale
Explore โActivation Steering
Activation Steering: Feature-Guided Interventions for Inference-Time Control
Explore โLong Context
Long Context Engineering: RoPE Scaling, KV Compression & Memory Optimization
Explore โSSMs & Hybrids
State Space Models & Hybrid Architectures: Mamba-2, Jamba, Griffin
Explore โMultimodal VLP
Multimodal Foundations: Vision Encoders, Contrastive Learning & Cross-Attention Fusion
Explore โTokens
Tokenization & Vocabulary Design
Explore โDecoding
Decoding & Sampling: Temperature, Top-p & Inference-Time Control
Explore โBackprop
Backpropagation & Automatic Differentiation
Explore โScore Matching
Score Matching & Score-Based Generative Models
Explore โICL
In-Context Learning: Learning Without Weight Updates
Explore โOT/Wasserstein
Optimal Transport & Wasserstein Distance
Explore โFlows
Normalizing Flows: Exact Likelihood via Invertible Transforms
Explore โPPO
PPO: Proximal Policy Optimization
Explore โResiduals
Residual Connections & Skip Connections
Explore โCFG
Classifier-Free Guidance in Diffusion
Explore โRAG
Retrieval-Augmented Generation (RAG)
Explore โAdversarial
Adversarial Examples & Robustness
Explore โGrokking
Grokking: Delayed Generalization
Explore โLogit Lens
Logit Lens: Probing Intermediate Representations
Explore โLR Schedules
Learning Rate Schedules: Warmup, Decay & Cycling
Explore โInit
Weight Initialization: Xavier, He & ยตP
Explore โContrastive
Contrastive Learning & InfoNCE
Explore โDistributed
Distributed Training: Data, Tensor & Pipeline Parallelism
Explore โBeam Search
Beam Search & Structured Decoding
Explore โDropout
Dropout: Stochastic Regularization
Explore โEBMs
Energy-Based Models & Score Functions
Explore โLayerNorm
Layer Normalization & RMSNorm
Explore โFisher Info
Fisher Information & Information Geometry
Explore โNatural Grad
Natural Gradient & Riemannian Optimization
Explore โSGD+Momentum
SGD & Momentum: The Workhorses of Optimization
Explore โAdamW
Weight Decay & AdamW: Decoupled Regularization
Explore โGrad Clip
Gradient Clipping & Explosion Prevention
Explore โLabel Smooth
Label Smoothing & Soft Targets
Explore โBatchNorm
Batch Normalization
Explore โDistillation
Knowledge Distillation: Learning from Teachers
Explore โQuantization
Quantization: Compressing Models to Integers
Explore โPruning
Pruning: Removing Unnecessary Weights
Explore โSSL
Self-Supervised Learning: Labels from Structure
Explore โCalibration
Calibration & Temperature Scaling
Explore โSwiGLU
SwiGLU & Gated Activations
Explore โFlashAttn
FlashAttention: IO-Aware Attention
Explore โConstitutional
Constitutional AI: Principles-Based Alignment
Explore โBregman
Bregman Divergence & Mirror Descent
Explore โRKHS
Reproducing Kernel Hilbert Spaces
Explore โTDA
Persistent Homology & Topological Data Analysis
Explore โLie Groups
Lie Groups & Equivariant Networks
Explore โTest-Time
Test-Time Compute & Inference Scaling
Explore โCoT
Chain-of-Thought Prompting
Explore โWorld Models
World Models & Model-Based RL
Explore โSynth Data
Synthetic Data & Self-Improvement
Explore โConsistency
Consistency Models: One-Step Diffusion
Explore โCheckpointing
Activation Checkpointing & Memory Efficiency
Explore โGQA
Grouped Query Attention (GQA)
Explore โPRMs
Process Reward Models
Explore โRLAIF
RLAIF: AI Feedback
Explore โFlow Match
Flow Matching & Rectified Flows
Explore โInstruct
Instruction Tuning
Explore โDeliberative
Deliberative Alignment
Explore โDebate
AI Safety via Debate
Explore โIDA
Iterated Amplification
Explore โWeakโStrong
Weak-to-Strong Generalization
Explore โAuto RedTeam
Automated Red Teaming
Explore โMesa-Opt
Mesa-Optimization & Inner Alignment
Explore โSleepers
Sleeper Agents & Alignment Faking
Explore โLLM-as-Judge
Model-Graded Evaluations
Explore โElicitation
Capability Elicitation & ELK
Explore โSandwich
Sandwiching Evaluations
Explore โMoD
Mixture-of-Depths
Explore โMCTS-LLM
Tree Search over Thoughts
Explore โVideoWM
Video World Models
Explore โSelf-Improve
Self-Improvement & Distillation Loops
Explore โCollapse
Model Collapse & Synthetic Data
Explore โInfCtx
Infinite Context Architectures
Explore โWhy These 100 Concepts?
Complete Coverage
Together, these concepts explain the core mechanisms behind language models, diffusion models, and multimodal systems.
Missing Intuition
Each concept includes what's still poorly explained in textbooks and papers - the intuition gaps we aim to fill.
Connected Knowledge
See how concepts build on each other. Understand prerequisites and what each idea unlocks.