Date:2026/7/17 14:30-16:00
Location:R110, CSIE
Speaker:Dr. Kangfu Mei, Google DeepMind Gemini Team
Host:Prof. Shao-Yuan Lo 羅紹元
Abstract:
Generative models, such as latent diffusion models, have transformed foundational image and video synthesis tasks, yet significant challenges remain in maximizing their generalization capabilities and real-world efficiency under constrained computational budgets. Shifting the paradigm from costly pre-training scaling to intelligent inference-time optimization, this talk first explores the inference scaling laws of text-to-image models, demonstrating that smaller models can sample more efficiently than larger ones when given an increased inference compute budget. To actively steer and enrich these generative priors at inference time, we introduce Kernel Density Steering (KDS)—a plug-and-play framework that applies patch-wise Mean Shift guidance to a sampled ensemble of latent particles to balance perceptual sharpness and fidelity—alongside a Multimodal Super-Resolution (MMSR) paradigm that sharpens the posterior distribution by combining text prompts with dense multimodal inputs. Finally, to bridge the gap between heavy computational demands and practical deployment, we present Conditional Diffusion Distillation (CoDi), which enforces a novel consistency score to reduce sampling costs by up to 99%, enabling high-fidelity conditional generation in as few as 1 to 4 steps. Together, these advancements outline a clear path toward more practical, scalable, and controllable multimodal generative AI.
Bio:
Kangfu Mei is a Senior Research Scientist at Google DeepMind Gemini Team. He was the core contributor of several foundational multimodal generative models, including Veo3, Genie3, and Gemini Omni. Before that, he got his Ph.D. degree from Johns Hopkins University, working on efficient and scalable multimodal generation.


