Shortcutting Time and Depth in Diffusion Models for Faster Agentic Applications

Date:

PDF Slides

This talk has been presented at Ai4 Conference.

Abstract: Diffusion models have emerged as a powerful alternative to conventional autoregressive LLMs, offering fast parallel token generation, built-in quality control, and native multi-modality support. However, their iterative denoising process remains computationally expensive for cost-effective deployments. This talk will detail recent few-step generation methods that substantially reduce sampling complexity. I will then bridge these technical optimization strategies with the broader architectural shift of integrating diffusion models into next-generation AI agents across language, video, audio, and robotics. Attendees will walk away with an understanding of recent trends in computationally efficient AI architectures, preparing them to build and deploy the next wave of diffusion-based VLAs (Vision-Language-Action Models), WAMs (World Action Models), and Computer Use Agents (CUAs).