根据任务阶段动态选择最优视角,提升精细操作成功率。
BFA: Best-Feature-Aware Fusion for Multi-View Fine-grained Manipulation
- 按任务阶段动态加权多视角特征,避免冗余信息
- 在多个任务上提升22%-46%成功率
- 可插拔设计,适配各类策略网络
现实场景中,多视角相机常用于精细操作任务。现有方法(如ACT)通常等权处理多视角特征并直接拼接,导致冗余视觉信息和更高计算开销,影响操作效果。针对精细操作常分多阶段且各阶段最优视角不同的特点,本文提出一种即插即用的最优特征感知融合(BFA)策略。基于策略网络的视觉主干,设计轻量级网络预测各视角重要性得分,再对多视角特征进行重加权融合,输入端到端策略网络,实现无缝集成。实验表明,该方法在不同任务上相比多个基线提升22%-46%的成功率,为解决精细操作中的关键挑战提供新思路。
原文摘要 · Abstract (English)
In real-world scenarios, multi-view cameras are typically employed for fine-grained manipulation tasks. Existing approaches (e.g., ACT) tend to treat multi-view features equally and directly concatenate them for policy learning. However, it will introduce redundant visual information and bring higher computational costs, leading to ineffective manipulation. For a fine-grained manipulation task, it tends to involve multiple stages while the most contributed view for different stages is varied over time. In this paper, we propose a plug-and-play best-feature-aware (BFA) fusion strategy for multi-view manipulation tasks, which is adaptable to various policies. Built upon the visual backbone of the policy network, we design a lightweight network to predict the importance score of each view. Based on the predicted importance scores, the reweighted multi-view features are subsequently fused and input into the end-to-end policy network, enabling seamless integration. Notably, our method demonstrates outstanding performance in fine-grained manipulations. Experimental results show that our approach outperforms multiple baselines by 22-46% success rate on different tasks. Our work provides new insights and inspiration for tackling key challenges in fine-grained manipulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。