Inception Labs introduces Mercury, the first commercial-scale diffusion large language model. https://www.inceptionlabs.ai/news
Why this is potentially disruptive from their release blurb: "Mercury Coder pushes the frontier of AI capabilities: it is 5-10x faster than the current generation of LLMs, providing high-quality responses at low costs. Our work builds on breakthrough research from our founders–who pioneered the first diffusion models for images—and who co-invented core generative AI techniques such as Direct Preference Optimization, Flash Attention, and Decision Transformers." Mercury is up to 10x faster than frontier speed-optimized LLMs. Our models run at over 1000 tokens/sec on NVIDIA H100s, a speed previously possible only using custom chips.
Implications for NVDA.