让机器人根据任务阶段自动调节视觉与触觉信息的权重,提升操作成功率。
Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation
- 基于力反馈动态调整视觉与触觉特征的融合权重,无需人工标注。
- 在三个高接触任务中实测成功率达93%,显著优于传统方法。
- 通过预测未来受力来增强触觉模态,适合复杂精细操作场景。
有效利用多模态数据对机器人泛化到多样任务至关重要。然而,不同模态间异质性使得融合困难。现有方法虽尝试全面融合特征,却常忽略各模态在不同操作阶段所需关注程度各异的事实。为此,我们提出一种力引导的注意力融合模块,可自适应调整视觉与触觉特征的权重,无需人工标注。同时引入自监督未来力预测辅助任务,强化触觉模态、缓解数据不平衡,并促进合理调整。该方法在真实世界实验中,于三个精细且高接触的任务上实现平均93%的成功率。进一步分析表明,策略能根据不同操作阶段适配地调整对各模态的关注度。视频演示见 https://adaptac-dex.github.io/。
原文摘要 · Abstract (English)
Effectively utilizing multi-sensory data is important for robots to generalize across diverse tasks. However, the heterogeneous nature of these modalities makes fusion challenging. Existing methods propose strategies to obtain comprehensively fused features but often ignore the fact that each modality requires different levels of attention at different manipulation stages. To address this, we propose a force-guided attention fusion module that adaptively adjusts the weights of visual and tactile features without human labeling. We also introduce a self-supervised future force prediction auxiliary task to reinforce the tactile modality, improve data imbalance, and encourage proper adjustment. Our method achieves an average success rate of 93% across three fine-grained, contactrich tasks in real-world experiments. Further analysis shows that our policy appropriately adjusts attention to each modality at different manipulation stages. The videos can be viewed at https://adaptac-dex.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。