EP 58: Every Millisecond Matters: Diffusion LLMs and the Future of Voice AI | Aditya Grover, Inception

August 23
15 mins

Episode Description

Recorded live at the Ai4 Podcast Pavilion, Sam wraps Day One with Aditya Grover, Co-Founder & CTO of Inception, on why the next generation of LLMs won't look anything like the ones we use today.

What's Covered:

"Every Millisecond Matters" — Why latency, not intelligence, is the real bottleneck holding back voice agents and multi-step AI agents alike.

How Mercury Actually Generates Text — Instead of predicting one token at a time like every autoregressive model, Mercury generates a rough draft of the full response and refines it into coherence — diffusion, applied to language instead of images.

Solving Voice AI's Impossible Tradeoff — Fast-but-lower-quality, or high-quality-but-too-slow: Aditya explains how Mercury 2 finally delivers both.

A Term Coined Live at This Conference — From Aditya's own Ai4 keynote: "We're moving from token maxing to value maxing."

Advice for the Next Generation — Ten-plus years into AI research, Aditya's honest take on why this is still the best time to pursue a PhD, join a startup, or do both.

The Next 5-10 Years of Voice AI — A prediction for a future where voice becomes humans' predominant mode of interacting with AI, the same way it is with each other.

Key Quote:

"Sequential generation is not a law of nature... AI can have a different way of generation, one that's more parallelizable."

Connect with Aditya:

LinkedIn: https://www.linkedin.com/in/aditya-grover/

Inception: https://www.inceptionlabs.ai/

Subscribe: Spotify | Apple Podcasts | Amazon Music | iHeart Radio | YouTube | Substack

#Ai4Conference #InceptionLabs #DiffusionLLM #VoiceAI #AsembleAI

See all episodes