AMD's MI355 offers superior cost and performance vs NVIDIA's B200; speed/cost advantage is durable moat, with further upside as MI355X catches CUDA parity on distributed inference.
MI355 is 40% cheaper than B200 on same architecture; distributed inference scaling provides cost-per-token advantage; execution speed is sustainable competitive moat
AMD ALERT 🚀 MI355 is now 40% cheaper than B200 on GLM5 architecture for Single Node serving FP8 14 weeks after the initial launch of GLM5 on both non-MTP & MTP with spec decode for SGLang v0.12 for both CUDA & ROCm. SPEED IS THE MOAT!! Great work to @AnushElangovan, @roaner, HaiShaw & his team! Next step is for MI355X to catch up to CUDA when composing production inference optimizations like FP4 & on distributed inferencing where you can gang up MI355 boxes such that per GPU performance goes up thus the cost per million tokens goes down.
| Horizon | Raw move | vs SPY |
|---|---|---|
| 1 day | +8.1% | +7.1% |
| 1 week | +21.7% | +19.4% |
| 1 month | +29.8% | +28.0% |
| 3 months | — | — |
| 6 months | — | — |