让视觉语言动作模型同时感知触觉和扭矩,提升真实世界交互能力
Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models

- 设计模块化感官流,分别处理触觉与扭矩信号并融合到动作预测中
- 在真实场景实验中,多信号联合使用使性能显著提升,实现协同增益
- 适合研究具身智能、机器人操作或多模态感知的开发者参考
人类理解并交互真实世界依赖于视觉以外的多种物理反馈。受此启发,现有方法尝试将物理感官信号引入视觉-语言-动作模型(VLAs),但通常仅关注单一类型信号,难以捕捉真实交互中多模态异构且互补的特性。本文提出MoSS——一种模块化感官流框架,可使VLAs有效利用多种感官信号进行动作预测。具体而言,通过解耦的模态流,采用联合跨模态自注意力机制将异构物理信号融入动作流;为确保新模态稳定融入,采用两阶段训练策略,初期冻结预训练VLA参数;此外,引入辅助任务以预测未来物理信号,更好地捕捉接触交互动态。大量真实世界实验表明,MoSS成功增强了VLAs对多样化物理信号(如触觉与扭矩)的利用能力,多信号融合带来协同性能提升。
原文摘要 · Abstract (English)
Humans understand and interact with the real world by relying on diverse physical feedback beyond visual perception. Motivated by this, recent approaches attempt to incorporate physical sensory signals into Vision-Language-Action models (VLAs). However, they typically focus on a single type of physical signal, failing to capture the heterogeneous and complementary nature of real-world interactions. In this paper, we propose MoSS, a modular sensory stream framework that adapts VLAs to leverage multiple sensory signals for action prediction. Specifically, we introduce decoupled modality streams that integrate heterogeneous physical signals into the action stream via joint cross-modal self-attention. To enable stable incorporation of new modalities, we adopt a two-stage training scheme that freezes pretrained VLA parameters in the early stage. Furthermore, to better capture contact interaction dynamics, we incorporate an auxiliary task that predicts future physical signals. Through extensive real-world experiments, we demonstrate that MoSS successfully augments VLAs to leverage diverse physical signals (i.e., tactile and torque), integrating multiple signals to achieve synergistic performance gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。