Google's TPUv7 is out! ML accelerator marketing material is usually pretty inscrutable (what numbers are even comparable?), so here I'll explain concretely how this compares with Nvidia. 🧵
Google TPUv7: - 4.6 PFLOP/s FP8 - 192 GB HBM @ 7.4 TB/s - 600 GB/s (unidi) ICI - ~1000 watts Nvidia GB200: - 5 PFLOP/s FP8 / 10 PFLOP/s FP4 - 192 GB HBM @ 8 TB/s - 900 GB/s (unidi) NVLink - ~1200 watts https://blog.google/products/g...
(disclaimer: everything here is just synthesis of public knowledge, and is my own opinion.) Roughly speaking, TPUv7 is about the same or slightly worse spec than GB200. It runs at slightly lower power, so you could probably consider them roughly even on perf/W. Jax / XLA will
In fact they're very similar in other ways - the package is almost identical. 8 stacks of HBM3e = 192 GB @ ~ 8 TB/s, surrounding two large compute dies. TPU appears to have moved I/O to a thin die at the top, which probably reduces the overall package cost a bit.
At the system level, TPU wins hard - ICI scales to 9,216 chips. However, the 3D torus topology limits programmability. Compare GB200: only 72 chips, but on a switched network. It's a much more flexible topology, but the switches consume power, and you have to lean on the
The blog post is quite hyperbolic and dishonest - making a comparison to El Capitan FP64 performance. The fair comparison there is against El Capitan FP8 peak perf, which is 43808 MI300A * 1961 TFLOP/s = 86 exaflops. This means that a TPUv7 9216-chip pod has about half the flops
One interesting tidbit is that this is likely what was supposed to be TPU v6p - a training chip. But (perhaps after reasoning models went big) it got renamed to TPU v7 and called "the first Google TPU for the age of inference" - quite the pivot :)
Overall it's a GB200-class chip within a superior scaffold: OCSes, racks, DCN, building design. Nothing revolutionary this time, just solid engineering. Great work by TPU team as always. Once again, Nvidia and Google are in a class of their own in the ML accelerator world. /🧵
@itsclivetime do you think it costs Google < $8k to make this? I know the memory is the expensive part
@itsclivetime Why wouldn't they sell this and grab a slice of the pie? Does GCP exclusivity + DeepMind's alpha really take in enough dough to offset that?
@snr_boost they are! see the bottom of the blog post
@itsclivetime This is v6p, not v7
@hassanience it used to be yes - but blog announces it as v7
@itsclivetime what's the 10x tflops/W graph being posted then?
@lennx_a50790 It's just straight 10x TFLOP/s, versus TPUv5e, which is two generations ago rather than one. This is also the first TPU to natively support FP8, so they're comparing against BF16 specs of past TPUs.
@itsclivetime The GB200 has 20 PFLOPs FP8, so about 10 PFLOPs per Blackwell Chip. And 40 PFLOPs of FP4 or 20 PFLOPs per chips. You way under cut the Blackwell chip. Blackwell is over twice the horsepower as the TPU7
@adrockdude Yes, these are their 4:2 sparse matmul "flops", which means they credit themselves with 2x flop throughput but have exactly the same number of floating point units on the chip as before. Also, to my knowledge almost nobody has publicly announced using Nvidia's sparse matmuls
@itsclivetime Thanks for writing this thread, very helpful
@itsclivetime great stuff, it seems for inference, a big win for Google right since they can sell a per hour rate at at least half of what a Blackwell would cost with similar performance and better power efficiency. Seems Google Cloud has a leg up on everyone now on cost efficiency per token?
@itsclivetime TPUv7 vs Nvidia—always a great debate. How does energy efficiency compare in real-world training loads?
@itsclivetime Who has cost advantage for a cloud service?
@itsclivetime Great breakdown! Always appreciate the comparison to Nvidia.
@itsclivetime That's a powerhouse for sure! Would be amazing if they offered a variant to consumers.
@itsclivetime @grok 用中文总结一下内容
@itsclivetime @grok Does this beat the V2 LPU of Groq?
@itsclivetime Nobody can buy a Google TPU, they only exist via Google cloud services, so I think you can never have a true apples-to-apples comparison to Nvidia.




