Just read rumors about the RTX 5090 sporting 28GB of VRAM. The number looks impressive, but let's do the math: with a 512-bit bus and GDDR7 at 28 Gbps, that's theoretically 1792 GB/s bandwidth.
The catch? For AI inference, bus width isn't everything—latency matters too. And if NVIDIA opts for HBM like Apple's M7 Ultra, thermal design becomes the true bottleneck.