Model & Hardware Sizing
Inputs
MoE decouples total parameter capacity from FLOPs consumed per forward pass.
70.0 B
Active weights stored in memory. Total weights $N_{\text{total}}$.
15.00 T
Total pre-training corpus tokens consumed during optimization.
FP8 doubles peak mathematical execution throughput relative to 16-bit.
16,384
Concurrent GPUs networked via NVLink / InfiniBand fabrics.
42 %
Efficiency after 3D tensor, pipeline, and data parallel communication overhead.
Blended cloud rental rate (Tier-1 provider with interconnect).
Total Pre-Training Compute
6.30e+24
6,300 ZettaFLOPs ($6 \times N_{\text{active}} \times D$)
Chinchilla Optimal Ratio
10.7 ×
Optimal: 1.40T Tokens ($D \approx 20N$)
Cluster Wall-Clock Time
53.8 Days
1,291 Total Continuous Hours
Estimated Training CapEx
$60.28 M
21.15 Million GPU-Hours
Empirical Loss vs. Compute Scaling Frontier
Hoffmann et al. (Chinchilla) Power-Law Loss Projection
Chinchilla Optimal
Selected Configuration
Pre-Training Token Allocation Strategy
Inference-Optimal Over-Trained (10.7×)
Active GPU Compute Power
13.62 EFLOPS
Peak theoretical: 32.43 ExaFLOPs
Predicted Test Perplexity
5.82
Estimated Loss: 1.761 nats
Inference Compute Amortization
Optimal
Downstream queries amortize training cost
Energy Footprint (MWh)
14,717 MWh
PUE ~1.2 with liquid-assisted cooling