Inception Labs is an AI research and product company building diffusion-based large language models (dLLMs) for production applications. Their Mercury family of models generates tokens in parallel rather than sequentially, delivering sub-300ms time-to-first-token, 5–7x higher throughput, and up to 70% lower cost per task compared to traditional autoregressive LLMs. The company is backed by leading venture capital firms and deploys its models at Fortune 500 companies, with the team drawn from Stanford, UCLA, Cornell, Google DeepMind, Meta AI, Microsoft AI, and OpenAI.