Inception Labs released a production-grade diffusion LLM
Diffusion Large Language Models (dLLMs) are revolutionizing text generation by refining entire sequences rather than predicting one token at a time. Mercury, the first commercial-scale dLLM from Inception Labs, generates over 1000 tokens per second on an NVIDIA H100, offering faster, more structured, and potentially smarter text generation compared to traditional models.