arXiv:2504.14573cs.ROcs.AI2025-04被引 2

用跨模态注意力筛选有效感官信息,提升机器人长序列操作能力

Modality Selection and Skill Segmentation via Cross-Modality Attention

  • 通过跨模态注意力动态选择关键感官输入
  • 实现从专家演示中自动分割基础技能动作
  • 适合研究多模态机器人控制与分层策略的学者

将触觉、音频等额外感官模态融入基础机器人模型面临维度灾难挑战。本文提出跨模态注意力(CMA)机制,在每个时间步识别并选择对动作生成最有效的模态。进一步将CMA应用于从专家演示中分割基础技能,并基于此构建分层策略,以解决长时序、高接触力的复杂操作任务。

原文摘要 · Abstract (English)

Incorporating additional sensory modalities such as tactile and audio into foundational robotic models poses significant challenges due to the curse of dimensionality. This work addresses this issue through modality selection. We propose a cross-modality attention (CMA) mechanism to identify and selectively utilize the modalities that are most informative for action generation at each timestep. Furthermore, we extend the application of CMA to segment primitive skills from expert demonstrations and leverage this segmentation to train a hierarchical policy capable of solving long-horizon, contact-rich manipulation tasks.

多模态感知分层控制机器人学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。