Posts  / #POST-231252
REDDIT

Why "Mixture of Experts" architecture is the ultimate bull case for memory demand

D
Jul 1, 2026 · 12:07

Wall Street is completely sleeping on how the latest shift in AI architecture is skewing the memory-to-compute ratio heavily in favor of RAM. Everyone thinks AI is all about Nvidia GPUs and pure math power, but the newest models are being built using an architecture called Mixture of Experts (MoE). Instead of being one giant brain that does massive math for every single word, an MoE model is broken up into hundreds of specialized "mini-brains." When you ask it a question, a digital traffic cop only wakes up the specific expert needed while the rest stay asleep. This keeps computing costs flat, which crucially allows tech companies to build and run exponentially bigger, smarter models than ever before.

But here is the massive catch that makes this the ultimate bull case for Micron: because this architecture unlocks these giant models, the physical size of the AI is exploding, causing the required memory capacity per unit of compute to skyrocket. Even though +90% of those experts are sleeping at any given second, the entire library of experts has to stay loaded into the high-bandwidth memory (HBM) 24/7 because the system never knows who it will need for the very next word. On top of that, if a GPU doesn’t have enough RAM capacity, you are forced to split the model across multiple chips. This triggers a massive communication bottleneck as data constantly flies back and forth between GPUs, severely tanking their efficiency and utilization rates. Buying chips with massive individual RAM capacity allows datacenters to keep the model localized, slashing that inter-GPU chatter and dramatically improving hardware utilization. We’ve entered a world where AI scaling isn't limited by how fast a chip can do math, but by physical VRAM capacity. If this AI buildout continues, the push for bigger MoE models means the demand for HBM and high-capacity DRAM is going to blow past what anyone has priced in. Long $MU.