Posts  / CLS  / #POST-214730
REDDIT

Should we actually care about the TPU vs GPU war? and why Celestica wins either way.

B
Dec 22, 2025 · 13:43

I keep seeing this debate framed like it’s VHS vs Betamax all over again. The TPU vs GPU debate matters if you’re trying to predict which chip company captures the most economics. It matters a lot less if you’re positioned in the infrastructure layer that grows either way. That’s the whole point of a picks-and-shovels bet. You don’t need to know which miner finds gold. You just need miners to keep digging.

Nvidia’s GPUs on one side. Google’s TPUs on the other. And the noise got louder after reports that Meta is in talks to spend billions using Google’s TPUs, potentially starting around 2027, and maybe renting TPU capacity sooner through Google Cloud.

So… should we care? Yes and no. Because Nvidia has been the default “engine” behind the AI boom. If a hyperscaler like Meta shifts meaningful workloads to TPUs, that’s a real signal. It says two things:

1. Companies want alternatives to Nvidia’s stack
2. The AI compute bill is so big that saving a few percent matters

Google also knows the real moat isn’t only the chip. It’s the software. Google has been working on making TPUs play nicer with PyTorch (the framework most people use), with Meta helping through a project described as “TorchTPU.”

That matters because the old TPU knock was simple: “Cool chip, but I don’t want to rewrite my whole stack.” Google is trying to remove that excuse.

GPUs and TPUs are good at different things: Training and inference don’t behave the same.

Training looks like building a brain from scratch. It’s messy. You try different architectures. You tune. You iterate. You break things. GPUs shine here because they’re flexible and the tooling is mature.

Inference looks like running a factory. Same model. Same tasks. Over and over. Speed matters. Power draw matters. Cost per answer matters.

That’s where Google built Ironwood to live. Google literally introduced Ironwood as a TPU designed specifically for inference at scale.

Under the hood, TPUs also go hard on the one thing neural nets eat for breakfast: matrix math. Google’s older TPU deep dive explains the core matrix unit uses a 256×256 systolic array, which is 65,536 ALUs doing multiply-and-add work in lockstep.

So no, this isn’t a simple “winner takes all” format war.

It’s more like trucks vs sports cars. Both move you but have different use cases.

Even if you’re convinced TPUs “win” long term, you still run into a physical reality as all this compute has to sit somewhere.

It needs:

* racks
* power delivery that doesn’t melt
* cooling that actually works at high density
* networking that can move insane amounts of data
* testing, integration, and logistics so it all shows up and runs

That’s the picks-and-shovels part of the AI boom.

And that’s why I keep coming back to Celestica (CLS).

Celestica doesn’t need to guess whether GPUs or TPUs “win.” They benefit from the build-out either way. You still need to assemble the systems, integrate racks, test them, and ship them. Celestica literally offers rack integration services across design, productization, testing, fulfillment, logistics, and lifecycle support.

On top of that, AI clusters create heat problems that air cooling can’t solve forever. Celestica talks openly about high-density AI racks needing advanced thermal management like liquid cooling.

So when I see headlines like “Meta might spend billions on TPUs,” I don’t just think “chip rivalry.”

I think: more racks, more networking, more cooling, more integration.

More shovels getting sold.