Generating each token requires sweeping the entire model weight tensor through GPU memory. At batch size 1, operational intensity is only ~1 FLOP/byte, causing the execution units to wait on HBM memory bandwidth.
| GPU Architecture | Tier | Bandwidth | VRAM | FP16 TFLOPS | Est. Decode Speed | Est. TTFT (1.5k) |
|---|