A kernel is not complete without a reference, a shape/dtype matrix, and clean sanitizer results.
Write the kernel.
Understand the system.
Prove it.
Eleven interactive atlases—from your first CUDA warp to vLLM serving and multi-GPU topologies—in one 12-week GPU Kernel Engineering application.
KERNELNSIGHTCUTLASSINFERENCENCCL
No performance claim without warm-up, quantiles, profiler evidence, and a controlled baseline.
The real goal is a portfolio-grade operator that works in PyTorch, compile, and serving workloads.
One application.
Eleven domains.
A route from foundations to capstone. Every atlas includes an interactive lab, a decision model, and an acceptance artifact.
Not a reading list.
A production system.
14–16 hours per week. Every week produces working code, correctness evidence, or a measurement report.
C++/Linux/CMake environment; inspect strides and layouts
Grid, block, warp, divergence, and the first safe kernel
HBM, shared memory, bank conflicts, and occupancy
torch.library, fake kernels, opcheck, and a first Triton kernel
Reference, stride-aware indexing, and CUDA/Triton twins
Activation + multiply fusion; register pressure
Stable reduction, masking, and online softmax
Scatter/update, races, and a broad test matrix
Warm-up, quantiles, roofline, and three profile studies
Tile policy and the first end-to-end fused kernel
vLLM, CUDA Graphs, NCCL, and communication cost
TTFT/ITL/throughput report, two 15%+ fusions, and defense
CUDA/Triton dual implementation
median gain across two fused kernels
Nsight studies
vLLM TTFT/ITL/ throughput report
interview defense