Home Blog Spatial Lab Disciplines Agentic Tools
Learn • AI Academy
IP Network Infrastructure About Connect
SYSTEMS ARCHITECTURE • PERFORMANCE ENGINEERING

LLM Inference Latency, Throughput & TTFT Systems Lab

Size your inference serving pipelines, analyze memory-bandwidth rooflines, evaluate Time-to-First-Token (TTFT) vs. inter-token decode latency, and simulate speculative decoding speedups across 10 GPU architectures.

10 GPU Profiles
H100, H200, B200, 4090, Apple M4
Roofline Model
Prefill vs Autoregressive Decode
Speculative ROI
Draft Model Tree Verification
Zero Fabrication
Hardware Bandwidth Physics
Estimated Total Response Latency
3.12 s
Single-Stream Decode: 88.4 tok/s • TTFT: 292 ms
Timing Distribution (TTFT vs Decode Duration) 9.4% Prefill / 90.6% Decode
Time to First Token (TTFT) 292 ms
Decode Phase Duration 2,828 ms
System Bottleneck: Memory-Bandwidth Bound (Autoregressive Decode)

Generating each token requires sweeping the entire model weight tensor through GPU memory. At batch size 1, operational intensity is only ~1 FLOP/byte, causing the execution units to wait on HBM memory bandwidth.

Inter-Token Latency (ITL)
11.3 ms
per generated token
Generation Speed
88.4 tok/s
single-stream throughput
Aggregate Throughput
88.4 tok/s
across 1 concurrent stream
Speculative Gain
2.28×
effective speedup multiplier
GPU Inference Hardware Comparison Matrix
Full Fleet Metrics
GPU Architecture Tier Bandwidth VRAM FP16 TFLOPS Est. Decode Speed Est. TTFT (1.5k)
Curriculum Integration

Theoretical Foundations & Architecture Guides

Master the engineering principles behind memory hierarchies, FlashAttention tiling, continuous batching, and speculative tree verification.