为视觉语言动作模型设计的自适应图像压缩方法,提升机器人控制性能。
Learned Image Compression for Vision-Language-Action Models

- 根据任务相关性动态分配比特率,结合时序上下文优化压缩
- 在相同码率下,控制成功率显著优于传统编码和现有学习压缩方法
- 适合带宽受限的远程机器人控制场景,实测效果更优
视觉语言动作(VLA)模型日益依赖高频多相机观测,导致视觉通信成为带宽受限或分布式部署场景下实时机器人控制的主要瓶颈。现有图像与视频编解码器侧重通用视觉保真度,而非下游VLA策略的控制性能。本文提出面向VLA机器人的学习图像压缩框架SPARC(SPatially Adaptive Rate Control)。核心观察是:视觉信息的重要性在不同摄像头视角及图像空间区域间差异显著。SPARC采用轻量级时序掩码选择器,基于任务相关性自适应分配潜在表示的码率,同时利用时序上下文。我们还引入倾斜率损失,通过减少基于熵的目标对罕见但关键视觉模式的过度抑制,稳定训练过程。在RoboCasa365、VLABench和LIBERO等多样化机器人基准测试中,SPARC在相同码率预算下持续优于传统编解码器和近期学习压缩方法。此外,在远程控制实际部署中,本方法显著改善了码率-成功效率权衡。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in bandwidth-constrained or distributed deployment settings. Existing image and video codecs, however, are designed to preserve generic visual fidelity rather than the control performance of downstream VLA policies. In this work, we introduce SPARC (SPatially Adaptive Rate Control), a learned image compression framework tailored for VLA-driven robots. Our key observation is that the importance of visual information varies substantially across both camera views and spatial regions within an image. Based on this observation, SPARC employs a lightweight temporal mask selector that adaptively allocates bitrate over latent representations according to task relevance while leveraging temporal context. We further introduce a tilted rate loss that stabilizes training by reducing the tendency of entropy-based objectives to over-suppress rare yet task-critical visual patterns. Experiments on diverse robotic benchmarks, including RoboCasa365, VLABench, and LIBERO, show that SPARC consistently achieves stronger control performance than conventional image/video codecs and recent learned compression methods under the same bitrate budget. We additionally demonstrate real-world deployment benefits in remote-control settings, where our method substantially improves the bitrate-success tradeoff.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。