Offline Semantic Guidance for Efficient Vision-Language-Action Policy Distillation
📰 ArXiv cs.AI
arXiv:2605.16241v1 Announce Type: cross Abstract: Billion-parameter Vision-Language-Action (VLA) policies have recently shown impressive performance in robotic manipulation, yet their size and inference cost remain major obstacles for real-time closed-loop control. We introduce \textbf{VLA-AD}, a distillation framework that uses a Vision-Language Model as an offline semantic supervisor to transfer large VLA teachers into lightweight student policies. Instead of relying only on low-level action i
DeepCamp AI