让机器人在6G网络下更高效执行语言指令,通过智能压缩视觉信息降低通信负担。
ComVLA: Communication-Aware Split Inference for VLA Models in 6G-Connected Robotics

- 用语言指令指导选择关键视觉特征,动态调整传输数据量。
- 仅传32个视觉令牌(原512个),推理计算降74%,延迟减22%。
- 适合6G环境下资源受限的移动机器人系统部署。
连接式机器人是6G新兴应用,移动机器人需根据自然语言指令操作物体。实现此功能的视觉-语言-动作(VLA)模型过大,无法在机器人本地运行,通常需将推理任务卸载至云端。然而无线链路限制了每控制周期可传输的感知数据量。现有方法中,语义通信编码器可压缩传感数据,但需针对特定信道重新训练;而VLA标记剪枝虽能减少图像标记数量,却忽略信道状况。本文提出ComVLA框架,利用语言中蕴含的密集语义信息来判断哪些视觉标记重要,从而根据信道容量自适应调整VLA标记预算。在LIBERO基准上,仅传输32个标记(相比原512个),相较OpenVLA-OFT基线,推理计算减少74%,延迟降低22%,任务成功率仅下降1.5个百分点(95.4% vs. 96.9%),且在瑞利和莱斯衰落信道下仍满足带宽约束。结果表明,协同设计VLA推理与无线通信是6G连接机器人的重要方向。
原文摘要 · Abstract (English)
Connected robotics is an emerging 6G application where mobile robots follow natural-language instructions to manipulate physical objects. The Vision-Language-Action (VLA) models that enable this are too large to run on the robot; a common trend is to offload inference to the cloud. The wireless link, however, limits how much sensing data the edge can transmit per control step. Two recent lines address this constraint: semantic communication codecs compress sensor data but require channel-specific retraining, and VLA token pruners select tokens from image but ignore the channel. Our insight is that the dense semantic information contained in the language already indicates which visual tokens matter. We propose ComVLA, a framework that uses this language guidance to adapt the VLA token budget to the channel capacity. Transmitting 32 tokens instead of 512 on the LIBERO benchmark, ComVLA cuts inference compute by 74% and inference latency by 22% versus the original OpenVLA-OFT baseline, at a cost of 1.5 pp in average task success (95.4% vs. 96.9%), and it stays within the capacity budget under Rayleigh and Rician fading. These results demonstrate that co-designing VLA inference and wireless communication is a practical direction for 6G-connected robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。