GPT-6 Astra is the start of the loop transformer era. AI hardware stock winners and losers.
**Disclaimer: All the paragraphs are written by a human. The table is generated by AI with supervision from a human.**
Enter, GPT-6 Astra. The Information [reports](https://www.theinformation.com/articles/secret-technique-behind-openais-astra-model-sparks-security-concerns) it uses a looped transformer architecture. It's a significant improvement over GPT5.6. This is the model that people say we've entered the AGI era.
LLMs have layers of data, called weights. During inference, **each layer gets computed** **once**. So moving weights from HBM memory to the chip is a major bottleneck because to generate a token, you need to move an entire layer into the chip, compute once, move it out, move the next layer in. Hence, memory bandwidth has been the #1 bottleneck for inference.
A loop transformer changes this. Instead of one layer being computed once, it will compute the **same layer multiple times**, increasing the model intelligence without relying memory bandwidth as much. Bandwidth is still important, but not as much anymore.
And because a looped transformer architecture can increase intelligence by running computing the same layer multiple times, you do not need to increase the model parameter count by as much to achieve the same level of intelligence. For example, a 250GB trillion parameter looped transformer model might be as smart as a 1 trillion model today.
In summary, loop transformer architecture decrease the bottleneck on memory bandwidth and capacity and increase the need for more compute.
Here are the winners and losers.
|**Hardware area**|**Loop Transformer impact**|**Winners**|**Losers / pressure**|
|:-|:-|:-|:-|
|**AI compute**|↑↑↑|NVIDIA, AMD, Google TPU, TSMC|Low-compute architectures|
|**HBM capacity**|↓ per FLOP|Compute vendors, hyperscalers|SK Hynix, Micron, Samsung|
|**HBM bandwidth**|↓ per FLOP|Compute vendors, hyperscalers|SK Hynix, Micron, Samsung|
|**On-chip SRAM / cache**|↑↑↑|TSMC, custom AI ASICs, SRAM IP suppliers, Samsung, Intel, Cerebras|HBM relatively less important|
|**GPU-to-GPU bandwidth**|↓|Hyperscalers|NVIDIA NVLink/NVSwitch, Broadcom/Marvell networking|
|**Dynamic / adaptive compute**|↑↑|NVIDIA CUDA, Google TPU/XLA, AMD ROCm|Less-flexible accelerator stacks|
|**Advanced packaging**|Shifts toward compute + cache|TSMC|HBM-heavy packaging relatively less important|
|**Total accelerator demand**|Probably ↑|All AI hardware vendors|Michael Burry, Reddit r/stocks|
|**Power + cooling**|↑|Vertiv, Eaton, Schneider|None obvious|
Also note, looped transformer architecture continues to prove The Bitter Lesson correct. [https://en.wikipedia.org/wiki/Bitter\_lesson](https://en.wikipedia.org/wiki/Bitter_lesson)