0/11 ATLASES
UNIFIED LEARNING SYSTEM · 2026

Write the kernel.
Understand the system.
Prove it.

Eleven interactive atlases—from your first CUDA warp to vLLM serving and multi-GPU topologies—in one 12-week GPU Kernel Engineering application.

View the 12 weeks
11unified atlases
12intensive weeks
5LLM operators
3evidence gates
LEARNING GRAPHONLINE
CUDATRITONMEMORYOPSGPU
KERNEL
NSIGHTCUTLASSINFERENCENCCL
ARCH sm_89TRACK 14–16 h/wMODE evidence-first
01
CORRECTNESS

A kernel is not complete without a reference, a shape/dtype matrix, and clean sanitizer results.

02
MEASUREMENT

No performance claim without warm-up, quantiles, profiler evidence, and a controlled baseline.

03
INTEGRATION

The real goal is a portfolio-grade operator that works in PyTorch, compile, and serving workloads.

ATLAS MAP

One application.
Eleven domains.

A route from foundations to capstone. Every atlas includes an interactive lab, a decision model, and an acceptance artifact.

12-WEEK INTENSIVE ROUTE

Not a reading list.
A production system.

14–16 hours per week. Every week produces working code, correctness evidence, or a measurement report.

01
Toolchain & tensor anatomy

C++/Linux/CMake environment; inspect strides and layouts

Foundation
02
CUDA mental model

Grid, block, warp, divergence, and the first safe kernel

CUDA
03
Memory & coalescing

HBM, shared memory, bank conflicts, and occupancy

Memory
04
PyTorch custom operator

torch.library, fake kernels, opcheck, and a first Triton kernel

Integration
05
RMSNorm & RoPE

Reference, stride-aware indexing, and CUDA/Triton twins

Operator
06
SwiGLU

Activation + multiply fusion; register pressure

Operator
07
Masked softmax & attention

Stable reduction, masking, and online softmax

Operator
08
KV-cache & correctness

Scatter/update, races, and a broad test matrix

Evidence
09
Benchmark & Nsight

Warm-up, quantiles, roofline, and three profile studies

Measurement
10
CUTLASS & fusion

Tile policy and the first end-to-end fused kernel

Optimization
11
Inference & multi-GPU

vLLM, CUDA Graphs, NCCL, and communication cost

Systems
12
Capstone & portfolio

TTFT/ITL/throughput report, two 15%+ fusions, and defense

Graduation
GRADUATION GATE

CUDA/Triton dual implementation

≥15%

median gain across two fused kernels

3

Nsight studies

1

vLLM TTFT/ITL/ throughput report

≥80%

interview defense