00
Start here
Three honest routes in. Pick the one that describes you now, not the one that describes who you want to be.
You have never trained anything Start at the bottom and do not skip the maths. Three to six months, honestly.
You can code and want to fine-tune this week Skip the theory for now. Run one fine-tune end to end, then come back for the parts that broke.
You want to understand a model, not just use one Write one. Everything after that reads differently.
01
Foundations 12
The linear algebra, calculus and probability every later section quietly assumes. Skip it if you already read a Jacobian without flinching.
Nothing in this section matches the filter.
02
Python & the tooling 11
You will spend more hours in NumPy, pandas and a terminal than in any model architecture.
Course CS50's Introduction to Programming with Python Harvard free · CC BY-NC-SA · from zero
Book Python Data Science Handbook Jake VanderPlas full text online · NumPy, pandas, matplotlib, scikit-learn
Docs NumPy: the absolute basics for beginners NumPy official, short, correct
Docs 10 minutes to pandas pandas the fastest honest introduction
Book Scientific Visualization: Python + Matplotlib Nicolas Rougier free PDF · how to make a figure that is not embarrassing
Course The Missing Semester of Your CS Education MIT shell, git, tmux, debugging · CC BY-NC-SA
Book Pro Git Chacon & Straub free · CC BY-NC-SA · the entire book online
Interactive Learn Git Branching Peter Cottle git as a visual puzzle game
Docs The Python Tutorial Python Software Foundation the official tour of the language, still the most accurate one
Docs uv — Python packaging that is not painful Astral the modern, fast way to manage environments
Docs Jupyter Notebook documentation Project Jupyter the tool every course below hands you
Nothing in this section matches the filter.
03
Classical machine learning 14
Gradient boosting still wins most tabular problems. Learn this before reaching for a transformer.
Course Machine Learning Crash Course Google free · ~15 h · interactive, rebuilt in 2024
Course Machine Learning Specialization Andrew Ng · DeepLearning.AI free to audit · the canonical entry point
Course Stanford CS229: Machine Learning Stanford lecture notes and problem sets, openly posted
Video CS229 lecture videos Stanford Online the full lecture series, Autumn 2018
Book An Introduction to Statistical Learning James, Witten, Hastie & Tibshirani free PDF · Python and R editions · plus a free video course
Book The Elements of Statistical Learning Hastie, Tibshirani & Friedman free PDF · the harder sequel
Course Kaggle Learn Kaggle ~20 micro-courses, a few hours each, all in-browser
Docs scikit-learn User Guide scikit-learn a textbook disguised as documentation
Course Machine Learning for Beginners Microsoft 12 weeks, 26 lessons · MIT licence
Course Data Science for Beginners Microsoft 10 weeks, 20 lessons · MIT licence
Course mlcourse.ai Yury Kashnitsky · OpenDataScience free · notebooks, assignments and Kaggle contests
Docs XGBoost tutorials XGBoost the algorithm that quietly wins competitions
Book Probabilistic Machine Learning Kevin Murphy free draft PDFs of both volumes
Docs Google Machine Learning Guides Google the rules-of-ML and tuning playbooks used internally
Nothing in this section matches the filter.
04
Deep learning 15
Backpropagation, optimisers, regularisation, convolutions — the layer of knowledge that does not go stale.
Course Practical Deep Learning for Coders fast.ai · Jeremy Howard free · top-down, code first · the best-loved course in the field
Course MIT 6.S191: Introduction to Deep Learning MIT free · one week of lectures and labs, re-recorded every year
Book Dive into Deep Learning Zhang, Lipton, Li & Smola free · maths, code and discussion on the same page
Book Deep Learning Goodfellow, Bengio & Courville free HTML · the reference text
Book Understanding Deep Learning Simon J.D. Prince free PDF, slides and notebooks · modern and beautifully drawn
Book Neural Networks and Deep Learning Michael Nielsen free · derives backprop by hand, gently
Course NYU Deep Learning (DS-GA 1008) Yann LeCun & Alfredo Canziani free · lectures, notes and notebooks
Course Stanford CS231n: Deep Learning for Computer Vision Stanford notes and assignments open · the CNN classic
Docs PyTorch Tutorials PyTorch official · start with 'Learn the Basics'
Course UvA Deep Learning Tutorials University of Amsterdam notebooks in both PyTorch and JAX
Course Deep Learning Specialization Andrew Ng · DeepLearning.AI free to audit · five courses
Code micrograd Andrej Karpathy MIT · ~150 lines · autograd small enough to read in one sitting
Book Deep Learning Tuning Playbook Google Research & Harvard how to actually choose hyperparameters, from people who do it daily
Code Annotated Paper Implementations labml.ai MIT · 60+ architectures implemented in PyTorch with the paper alongside
Article A Recipe for Training Neural Networks Andrej Karpathy the debugging discipline nobody teaches you
Nothing in this section matches the filter.
05
NLP & transformers 14
Tokenisation, embeddings, attention — where language models actually begin.
Course Stanford CS224n: NLP with Deep Learning Christopher Manning · Stanford slides, notes and assignments open every year
Video CS224n lecture videos Stanford Online the full lecture series on YouTube
Course Hugging Face NLP Course Hugging Face free · Apache-2.0 · transformers, datasets and tokenizers, hands-on
Article The Illustrated Transformer Jay Alammar the diagram everyone has seen, and for good reason
Code The Annotated Transformer Harvard NLP the 2017 paper, line by line, as runnable code
Book Speech and Language Processing (3rd ed. draft) Jurafsky & Martin free draft chapters · the field's standard textbook
Course CMU CS11-711: Advanced NLP Graham Neubig · CMU slides, videos and assignments, all public
Article The Illustrated Word2vec Jay Alammar embeddings from first principles
Article Transformers from Scratch Brandon Rohrer a very patient derivation, matrices included
Article The Transformer Family v2.0 Lilian Weng every architectural variant, catalogued
Article Attention? Attention! Lilian Weng where attention came from
Paper Attention Is All You Need Vaswani et al. the 2017 paper itself · read it after the illustrated version
Course Advanced NLP with spaCy Ines Montani free · interactive · the practical pipeline side
Code NLP with Transformers — notebooks Tunstall, von Werra & Wolf the book's code, Apache-2.0, free to run
Nothing in this section matches the filter.
06
Build an LLM from scratch 16
The fastest way to stop treating a model as a black box is to write one. Everything here ends in code you typed yourself.
Course Neural Networks: Zero to Hero Andrej Karpathy free · ~10 video lectures · micrograd, then makemore, then GPT
Video Let's build GPT: from scratch, in code, spelled out Andrej Karpathy 2 h · a working GPT in one sitting
Video Let's build the GPT Tokenizer Andrej Karpathy 2 h 13 · BPE, and why tokenisation causes half of all LLM weirdness
Video Deep Dive into LLMs like ChatGPT Andrej Karpathy 3 h 31 · pretraining, SFT and RLHF end to end, little maths required
Code nanoGPT Andrej Karpathy MIT · ~300 lines that reproduce GPT-2
Code minGPT Andrej Karpathy MIT · the readable, minimal ancestor of nanoGPT
Code llm.c Andrej Karpathy MIT · GPT-2 training in plain C and CUDA, no framework
Code nanochat Andrej Karpathy the full ChatGPT pipeline — pretrain, SFT, RL, web UI — in one clean repo
Course Stanford CS336: Language Modeling from Scratch Stanford the whole pipeline as coursework: tokenizer, model, training, evaluation
Video CS336 lecture videos Stanford Online all lectures, free
Code Build a Large Language Model (From Scratch) — code Sebastian Raschka Apache-2.0 · every chapter's notebooks are free; the book is optional
Course LLM Course Maxime Labonne Apache-2.0 · roadmaps plus dozens of free Colab notebooks
Course Hugging Face LLM Course Hugging Face the NLP course's successor: pretraining, fine-tuning, inference
Code modded-nanogpt Keller Jordan et al. the community speedrun — read the diffs to see what actually helps
Code litgpt Lightning AI Apache-2.0 · 20+ models as readable single-file implementations
Article The Annotated GPT-2 Aman Arora GPT-2 explained alongside its own code
Nothing in this section matches the filter.
07
Pretraining & scale 14
What changes when the run stops fitting on one GPU: parallelism, schedules, scaling laws, and the failure modes nobody warns you about.
Book The Ultra-Scale Playbook: Training LLMs on GPU Clusters Hugging Face free · 5D parallelism, ZeRO, kernels · drawn from 4000+ experiments
Book The Smol Training Playbook Hugging Face how a small model actually gets trained, decision by decision
Code EleutherAI Cookbook EleutherAI MIT · practical recipes and calculators from people who train at scale
Article Transformer Math 101 EleutherAI the arithmetic of memory, FLOPs and cost, before you rent a GPU
Book How to Scale Your Model Google DeepMind free · a systems view of TPUs, parallelism and rooflines
Paper Scaling Laws for Neural Language Models Kaplan et al. the original scaling-law paper
Paper Training Compute-Optimal LLMs (Chinchilla) Hoffmann et al. the correction that reset the field's data budgets
Code OLMo — a fully open language model Allen Institute for AI weights, data, code and training logs, all released
Code nanotron Hugging Face Apache-2.0 · minimalist 3D-parallel pretraining
Code Megatron-LM NVIDIA the reference implementation for large-scale training
Docs DeepSpeed tutorials Microsoft ZeRO, offload and pipeline parallelism, explained by its authors
Docs Getting Started with FSDP PyTorch the sharding you will actually reach for first
Docs Hugging Face Accelerate Hugging Face the same training loop on one GPU, eight GPUs or a TPU
Paper OPT-175B logbook Meta AI the unvarnished day-by-day diary of a large run going wrong
Nothing in this section matches the filter.
08
Fine-tuning 28
Taking a trained model and bending it to your task. The largest section here, because it is the part most people actually need and the part with the most folklore around it.
Docs Fine-tune a pretrained model Hugging Face Transformers official · the honest starting point, Trainer and a plain PyTorch loop
Docs PEFT documentation Hugging Face LoRA, prefix tuning, IA3, adapters · conceptual guides plus API
Docs TRL — Transformer Reinforcement Learning Hugging Face the library behind most open SFT, DPO and GRPO recipes
Docs SFTTrainer guide Hugging Face TRL supervised fine-tuning, including packing and completion-only loss
Docs Unsloth documentation Unsloth free · fine-tuning that fits in a free Colab GPU, with the tricks explained
Notebook Unsloth notebooks Unsloth 100+ ready Colab notebooks — Llama, Qwen, Gemma, Mistral, vision, GRPO
Docs Axolotl documentation Axolotl AI Apache-2.0 · config-file fine-tuning; the YAML examples are the real course
Code LLaMA-Factory hiyouga et al. Apache-2.0 · 100+ models, a web UI, and every method in one place
Docs torchtune PyTorch native PyTorch recipes you can read top to bottom
Paper LoRA: Low-Rank Adaptation of Large Language Models Hu et al. the paper that made fine-tuning affordable
Paper QLoRA: Efficient Finetuning of Quantized LLMs Dettmers et al. 65B on a single 48 GB card · read alongside the bitsandbytes docs
Article Making LLMs even more accessible with bitsandbytes and QLoRA Hugging Face the practical write-up of 4-bit fine-tuning
Article Parameter-Efficient Fine-Tuning using PEFT Hugging Face the short version, with working code
Code Alignment Handbook Hugging Face Apache-2.0 · full, reproducible recipes: SFT then DPO, configs included
Course smol-course Hugging Face a hands-on course on aligning small models on modest hardware
Course Finetuning Large Language Models DeepLearning.AI & Lamini free short course · ~1 h · when to fine-tune at all
Article Practical Tips for Finetuning LLMs Using LoRA Sebastian Raschka hundreds of ablations distilled into rules of thumb
Article Finetuning LLMs Efficiently with Adapters Sebastian Raschka what each PEFT method actually changes in the weights
Docs Chat templates Hugging Face Transformers the single most common cause of a fine-tune that trains but never answers
Article Fine-tune Llama 3 with ORPO Maxime Labonne SFT and preference alignment collapsed into one step
Notebook Llama Cookbook Meta official fine-tuning and deployment recipes
Notebook Gemma Cookbook Google the same idea for Gemma, including Keras and JAX paths
Paper Finetuned Language Models Are Zero-Shot Learners (FLAN) Wei et al. the paper that introduced instruction tuning
Paper Self-Instruct Wang et al. how to generate the instruction data you do not have
Paper LIMA: Less Is More for Alignment Zhou et al. 1000 carefully chosen examples beating far larger sets
Code mergekit Arcee AI merging fine-tuned checkpoints instead of retraining
Notebook Hugging Face Open-Source AI Cookbook Hugging Face dozens of task-shaped recipes, all runnable
Notebook PyTorch Lightning / Fabric fine-tuning studios Lightning AI free-tier GPU studios with published fine-tuning templates
Nothing in this section matches the filter.
09
Preference tuning & RLHF 14
Turning a model that completes text into one that answers. RLHF, DPO, GRPO, and the reasoning-model recipes that followed.
Nothing in this section matches the filter.
10
Reinforcement learning 8
The background the section above assumes. Worth a detour if policy gradients are still a rumour to you.
Nothing in this section matches the filter.
11
Evaluation 10
The part people skip, and the reason so many fine-tunes look great in a demo and fail in use.
Nothing in this section matches the filter.
12
Data & datasets 11
Model quality is mostly data quality. This is the least glamorous section and the highest-leverage one.
Nothing in this section matches the filter.
13
Quantization, inference & speed 15
Getting a model to fit, and then to be fast. GPU arithmetic, kernels, quantization formats and serving.
Docs Quantization Hugging Face Transformers every supported format compared in one table
Docs bitsandbytes bitsandbytes 8-bit and 4-bit, the backend behind QLoRA
Code llama.cpp ggml.org MIT · GGUF, CPU inference, and the quant formats everyone ships
Docs vLLM documentation vLLM Apache-2.0 · paged attention and continuous batching, explained by the implementers
Article Making Deep Learning Go Brrrr From First Principles Horace He compute-bound, memory-bound, overhead-bound — the mental model
Article Optimizing LLMs for speed and memory Hugging Face quantization, Flash Attention and KV caching, with the arithmetic shown
Paper FlashAttention Dao et al. the IO-aware attention kernel now in everything
Paper GPTQ Frantar et al. post-training quantization down to 3-4 bits
Paper AWQ: Activation-aware Weight Quantization Lin et al. the other format you will meet on the Hub
Course GPU MODE lectures GPU MODE free · a whole community-run course on CUDA and kernel writing
Docs Triton tutorials OpenAI writing GPU kernels in Python, starting from vector add
Docs CUDA C++ Programming Guide NVIDIA the primary source, when the tutorials run out
Docs Optimum Hugging Face ONNX, TensorRT and hardware-specific export paths
Code Ollama Ollama MIT · the shortest path to running your fine-tune locally
Docs Text Generation Inference Hugging Face production serving, with the trade-offs documented
Nothing in this section matches the filter.
14
RAG, embeddings & retrieval 10
Putting the knowledge in the context instead of in the weights. Usually cheaper and more correctable than fine-tuning.
Nothing in this section matches the filter.
15
Prompting & agents 12
The cheapest lever, and the one to exhaust before you fine-tune anything.
Nothing in this section matches the filter.
16
Vision, audio & diffusion 12
Everything that is not text: image generation, vision-language models, speech.
Nothing in this section matches the filter.
17
Interpretability & safety 10
What is actually happening inside the weights, and what to do about the parts you would rather were not.
Nothing in this section matches the filter.
18
Shipping it: MLOps & systems 10
The distance between a notebook that works and a service that keeps working.
Nothing in this section matches the filter.
19
Staying current 10
The field moves fast enough that a course from last year has gaps. These are the feeds worth keeping.
Reference Daily Papers Hugging Face a curated, discussed selection rather than the full arXiv firehose
Reference arXiv cs.CL — recent arXiv the firehose itself, when you want it
Blog Lil'Log Lilian Weng long, careful surveys · several are the best sources on their topic
Blog Ahead of AI Sebastian Raschka research summaries with working code, free archive
Reference Connected Papers Connected Papers free tier · the citation graph around a paper, as a map
Reference alphaXiv alphaXiv arXiv with a comment layer, useful for hard papers
Blog The Gradient The Gradient longer-form essays, editorially reviewed
Blog Import AI Jack Clark weekly · research plus policy, from someone who reads everything
Reference AI Index Report Stanford HAI free · the annual numbers, with sources
Article How to Read a Paper S. Keshav three pages · the three-pass method, and it works
Nothing in this section matches the filter.
20
Free compute & practice 8
Nothing on this page sticks until you run it. All of these give you a GPU without a credit card.
Nothing in this section matches the filter.
Nothing on this page is hosted here and nothing has been copied from it: every row links
to its source, and the one-line notes are ours. Licences belong to the original authors —
several items are Apache-2.0, MIT or CC BY-NC-SA, and that is stated where it matters.
Found something dead, or something missing that should be here?
Tell us .
Sane Labs · 2026 ·
the lab ·
Synth-2
↑ Back to top