融合本体与外部传感器,提升机器人动作分割精度。
M2R2: MultiModal Robotic Representation for Temporal Action Segmentation
- 设计多模态特征提取器,融合本体与视觉信息
- 在三个机器人数据集上达到新最佳性能
- 支持特征复用,适合多任务机器人系统
时间动作分割(TAS)是机器人与计算机视觉领域的关键研究方向。机器人领域传统依赖本体感知信息识别技能边界,近年手术机器人开始引入视觉信息;而计算机视觉则主要依赖摄像头等外部传感器。现有机器人多模态TAS模型将特征融合嵌入模型内部,难以跨模型复用已学特征。同时,计算机视觉中常用的预训练视觉特征提取器在物体可见性受限场景下表现不佳。本文提出M2R2,一种专为TAS设计的多模态特征提取器,融合本体与外部传感器信息,并引入新型训练策略,实现特征在多个TAS模型间的复用。在三个机器人数据集REASSEMBLE、(Im)PerfectPour和JIGSAWS上取得新最优结果。通过详尽消融实验,验证了不同模态在机器人TAS任务中的贡献。
原文摘要 · Abstract (English)
Temporal action segmentation (TAS) has long been a key area of research in both robotics and computer vision. In robotics, algorithms have primarily focused on leveraging proprioceptive information to determine skill boundaries, with recent approaches in surgical robotics incorporating vision. In contrast, computer vision typically relies on exteroceptive sensors, such as cameras. Existing multimodal TAS models in robotics integrate feature fusion within the model, making it difficult to reuse learned features across different models. Meanwhile, pretrained vision-only feature extractors commonly used in computer vision struggle in scenarios with limited object visibility. In this work, we address these challenges by proposing M2R2, a multimodal feature extractor tailored for TAS, which combines information from both proprioceptive and exteroceptive sensors. We introduce a novel training strategy that enables the reuse of learned features across multiple TAS models. Our method sets a new state-of-the-art performance on three robotic datasets REASSEMBLE, (Im)PerfectPour, and JIGSAWS. Additionally, we conduct an extensive ablation study to evaluate the contribution of different modalities in robotic TAS tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。