Shrinking Vision Language Models for Edge Deployment via Distillation and Pruning
Date:
This talk has been presented at Embedded Vision Summit.
Abstract: The next generation of AI agents is moving beyond cloud-based text-only models and will interact with the physical world in real-time. However, deploying massive Vision-Language Models (VLMs) with billions of parameters on embedded devices remains a significant engineering hurdle. Drawing on our recent ICML and CVPR research papers, this session explores advancements in VLM optimization, specifically how distillation and pruning transform ‘heavyweight’ models into lean, edge-ready engines. I will analyze the trade-offs between model size and accuracy, providing practical insights for building responsive, multimodal agents that function efficiently without a persistent cloud connection.
