统一建模视觉触觉序列,提升机器人操作的跨模态感知能力
VT-MUSE: Multimodal Unified Sequential Visuotactile Representation Learning for Manipulation

- 分两阶段学习视觉触觉序列表示,强化跨模态时间对齐
- 在仿真和真实场景中均比最强基线提升11个百分点
- 适合需要精细触觉反馈的机器人抓取与操作任务
我们提出VT-MUSE,一种面向视觉触觉操作的多模态统一序列表征学习框架。现有方法常独立编码视觉与触觉观测后再融合,难以捕捉细粒度跨模态依赖;且多数方法仅关注当前时刻观测,忽视接触的时序演化。VT-MUSE通过双阶段表征学习解决上述问题:第一阶段,通过跨模态时间对齐与掩码视图一致性,联合优化各模态专用编码器;第二阶段,基于条件变分潜空间模型,处理掩码视觉序列与完整触觉历史,辅助解码器重建近期视觉观测并预测触觉深度变化,促使潜表示同时保留全局视觉上下文与局部接触动态。所学表示通过门控交叉注意力集成至轻量Transformer策略网络。在仿真基准测试中,VT-MUSE在所有任务上均优于最强基线11个百分点,并在真实世界实验中也取得显著提升。
原文摘要 · Abstract (English)
We propose VT-MUSE, a Multimodal Unified SEquential representation learning framework for visuotactilemanipulation. Existing approaches often encode visual and tactile observations independently before fusion, limiting their ability to capture fine-grained cross-modal dependencies. Moreover, most methods focus on observations at the current time step and overlook the temporal evolution of contact. VT-MUSE addresses both limitations through a two-stage representation learning framework. In Stage I, modality specific encoders are jointly adapted via cross-modal temporal alignment and masked-view consistency. In Stage II, a conditional variational latent model processes masked visual sequences together with full tactile histories. Auxiliary decoders reconstruct the masked recent visual observations and predict tactile depth changes, encouraging the latent representation to retain both global visual context and local contact dynamics. The learned representation is subsequently integrated into a lightweight Transformer policy through gated cross-attention. On the simulation benchmark, VT-MUSE outperforms the strongest baseline evaluated on all tasks by 11 percentage points and also achieves substantial improvements in real-world experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。