Compilers & Systems
Machine Learning Compilers
Bridging Models and Hardware for Maximum Performance
by Laxmen Murali · First edition, 2026
A first-principles path that takes you from a single transistor to writing production GPU kernels, and then to building an ML compiler of your own.
- Format
- Pages
- 1,068
- Chapters
- 63
- Parts
- 9
- Delivered by email in about a minute
- Free updates to this edition, forever
- Payment secured by Razorpay
- Read on any device that opens a PDF
About this book
Every time a model runs fast, a compiler made a decision. It fused two operators so a tensor never touched DRAM. It picked a tile size that kept the L2 warm. It chose bf16 over fp32 and proved the numerics still held. Those decisions are where most of the performance in modern AI comes from, and almost nobody outside a handful of teams understands how they are made.
This book is the missing curriculum. It starts at the switch, a physical thing that is either on or off, and does not stop until you have designed an intermediate representation, written the optimisation passes that walk it, generated code for a GPU, and benchmarked the result against PyTorch. There is no chapter where you are asked to accept something on faith.
It is deliberately long. A book this size is not padding; it is the honest measure of the gap between "I can train a model" and "I can explain, in cycles, why this kernel falls short of what the roofline says it should achieve." Every chapter carries the same nine sections: a story, the theory, the mathematics, the engineering reality, a performance analysis, a build-it exercise in real code, a debugging guide, interview questions, and the misconceptions that trip people up.
Who it's for
- ML engineers who are tired of `torch.compile` being a black box
- Compiler engineers moving from classical backends into the AI stack
- Systems and performance engineers targeting GPUs and accelerators
- Students who want the whole ladder, not one rung of it
- Anyone preparing for ML-compiler, CUDA, or performance-engineering interviews
Prerequisites
You should be able to write a program in any language. That's it. Python and C++ are taught from scratch where the book needs them, and no prior compilers, CUDA, or machine-learning background is assumed.
Genuinely first principles
Chapter 1 begins with a switch and ends with you having built a working CPU simulator, TOY-100, in Python. Every abstraction above it is constructed, not assumed.
The whole stack, one narrative
Logic gates, caches, C++, SSA, LLVM, MLIR, tensors, transformers, CUDA, fusion, Triton and FlashAttention. One continuous thread, not nine disconnected books.
You build things, constantly
A cache simulator in 25 lines. malloc-lite. A bytecode interpreter running real CPython opcodes. A register allocator. And finally, a complete ML compiler.
Performance treated as a science
Roofline, Amdahl, Gustafson, Little's Law, tail percentiles, and the seven ways benchmarks lie, applied to measured numbers rather than hand-waved ones.
Real stacks, dissected
TVM, PyTorch Dynamo and Inductor, XLA, IREE, Triton, TensorRT and ONNX Runtime. What each one actually does differently, and why.
Written for the interview and the job
Every chapter ends with the questions that get asked at compiler teams, the mistakes that get made in review, and the debugging playbook for when it breaks in production.
Full table of contents
Part I
Foundations of Computing
- 1 What Is a Computer? From Switches to Thinking Machines
- 2 Binary: The Language of Machines
- 3 Memory: The Computer’s Notebook
- 4 The CPU: The Engine of Computation
- 5 Caches and the Memory Hierarchy
- 6 Processes, Threads, and Virtual Memory
- 7 Performance: Thinking Like a Speed Engineer
Part II
Programming and Systems
- 8 What Is Programming? Variables, Types, and Functions
- 9 Pointers, References, and the Shape of Memory
- 10 Recursion and the Call Stack
- 11 Data Structures: Organizing Information
- 12 Algorithms and Complexity
- 13 Python: How the Interpreter Really Works
- 14 C++ for Systems and Compiler Engineers
Part III
Compiler Fundamentals
- 15 What Is a Compiler? The Grand Tour
- 16 Lexing: Turning Characters into Words
- 17 Parsing: Turning Words into Structure
- 18 ASTs and Semantic Analysis
- 19 Intermediate Representations and SSA
- 20 Control Flow Graphs and Data Flow Analysis
- 21 Classical Optimizations
- 22 The Backend: Instruction Selection, Scheduling, Register Allocation, Codegen
- 23 LLVM: The Industry Workhorse
- 24 MLIR: The Compiler Construction Kit
Part IV
Mathematics and Machine Learning
- 25 Vectors and Matrices from First Principles
- 26 Tensors, Broadcasting, and Memory Layout
- 27 Numerical Precision: FP32, FP16, BF16, FP8, INT8
- 28 What Is Machine Learning?
- 29 Neural Networks and Backpropagation
- 30 Convolutional Neural Networks
- 31 Transformers and Attention
- 32 Training, Inference, and Large Language Models
Part V
GPUs and Parallel Hardware
- 33 Why GPUs? The Parallel Revolution
- 34 CUDA and the SIMT Programming Model
- 35 The GPU Memory Hierarchy
- 36 Tensor Cores, Occupancy, and Kernel Scheduling
Part VI
Inside an ML Compiler
- 37 Anatomy of an ML Compiler: Graph IR and Tensor IR
- 38 Shape Inference and Graph Optimization
- 39 Operator and Kernel Fusion
- 40 Scheduling: Separating What from How
- 41 Memory Planning
- 42 Cost Models and Auto-Tuning
- 43 Kernel Generation
Part VII
Production Compiler Stacks
- 44 TVM: The Full-Stack Pioneer
- 45 PyTorch Compilation: FX, Dynamo, and Inductor
- 46 XLA and OpenXLA
- 47 IREE and the MLIR-Based Stacks
- 48 Triton: Python-to-GPU Kernels
- 49 TensorRT and ONNX Runtime
Part VIII
High-Performance Kernels
- 50 GEMM: From Naive to Near-Peak
- 51 Convolution Kernels
- 52 Attention Kernels and FlashAttention
- 53 Normalization, Activations, and Reductions
- 54 Advanced Loop Craft
- 55 Inference Serving: KV Cache, Continuous Batching, and Speculative Decoding
- 56 Parallelism: Tensor, Pipeline, and Expert
- 57 Quantization
Part IX
Build Your Own ML Compiler
- 58 Design and Architecture
- 59 Building the IR and Optimization Passes
- 60 The Backend and Code Generation
- 61 Testing, Benchmarking, and Debugging
- 62 How Real Teams Build ML Compilers
- 63 The Road Ahead: Careers, Interviews, and the Future