EP 58: Every Millisecond Matters: Diffusion LLMs and the Future of Voice AI | Aditya Grover, Inception
Episode Description
Recorded live at the Ai4 Podcast Pavilion, Sam wraps Day One with Aditya Grover, Co-Founder & CTO of Inception, on why the next generation of LLMs won't look anything like the ones we use today.
What's Covered:
"Every Millisecond Matters" — Why latency, not intelligence, is the real bottleneck holding back voice agents and multi-step AI agents alike.
How Mercury Actually Generates Text — Instead of predicting one token at a time like every autoregressive model, Mercury generates a rough draft of the full response and refines it into coherence — diffusion, applied to language instead of images.
Solving Voice AI's Impossible Tradeoff — Fast-but-lower-quality, or high-quality-but-too-slow: Aditya explains how Mercury 2 finally delivers both.
A Term Coined Live at This Conference — From Aditya's own Ai4 keynote: "We're moving from token maxing to value maxing."
Advice for the Next Generation — Ten-plus years into AI research, Aditya's honest take on why this is still the best time to pursue a PhD, join a startup, or do both.
The Next 5-10 Years of Voice AI — A prediction for a future where voice becomes humans' predominant mode of interacting with AI, the same way it is with each other.
Key Quote:
"Sequential generation is not a law of nature... AI can have a different way of generation, one that's more parallelizable."
Connect with Aditya:
LinkedIn: https://www.linkedin.com/in/aditya-grover/
Inception: https://www.inceptionlabs.ai/
Subscribe: Spotify | Apple Podcasts | Amazon Music | iHeart Radio | YouTube | Substack
#Ai4Conference #InceptionLabs #DiffusionLLM #VoiceAI #AsembleAI