You've walked the silicon end to end: what an SM and a warp are and why occupancy decides your fraction of peak, how the memory pyramid puts HBM bandwidth at the floor, and how the generations — Hopper, Ada, Blackwell, and what Rubin promises — trade compute for memory for cost. You've seen how Grace lends the GPU coherent host memory, how instances map to your bill, and where the non-NVIDIA shelf wins. Prove it stuck before you take it local.
Eight questions. Score 80% or better to unlock the local-inference finale. Every option explains itself, and you can take another pass any time.
The hardware core — compute, memory, and the shelf.
SMs, warps, and occupancy; the memory hierarchy and HBM ceiling; the GPU generations and their trade-offs; Grace host memory; instances; and the accelerators beyond NVIDIA.