让机器人一次计算最佳视角,提升抓取成功率。
A General One-Shot Multimodal Active Perception Framework for Robotic Manipulation: Learning to Predict Optimal Viewpoint
- 一次推理预测最优相机位置,摆脱反复试错
- 真实场景下抓取成功率接近翻倍,无需调参
- 适配不同任务,可直接从仿真迁移到现实
基于视觉的机器人操作中的主动感知旨在移动摄像头至更富信息量的观察视角,从而为下游任务提供高质量感知输入。现有方法多依赖迭代优化,导致时间和动作成本高,且与特定任务强耦合,限制了泛化能力。本文提出一种通用的一次性多模态主动感知框架,可直接推断最优视角,包含数据采集流程和最优视角预测网络。该框架将视角质量评估与整体结构解耦,支持异构任务需求。通过系统采样与评估候选视角构建大规模训练数据集,并采用领域随机化生成多样化样本。同时设计多模态视角预测网络,利用交叉注意力对齐融合多模态特征,直接预测相机位姿调整。该框架在视角受限环境下的机器人抓取中实现实例化。实验表明,基于该框架的主动感知显著提升抓取成功率。值得注意的是,真实世界评估中抓取成功率近乎翻倍,且实现无缝的仿真到现实迁移,无需额外微调,验证了框架的有效性。
原文摘要 · Abstract (English)
Active perception in vision-based robotic manipulation aims to move the camera toward more informative observation viewpoints, thereby providing high-quality perceptual inputs for downstream tasks. Most existing active perception methods rely on iterative optimization, leading to high time and motion costs, and are tightly coupled with task-specific objectives, which limits their transferability. In this paper, we propose a general one-shot multimodal active perception framework for robotic manipulation. The framework enables direct inference of optimal viewpoints and comprises a data collection pipeline and an optimal viewpoint prediction network. Specifically, the framework decouples viewpoint quality evaluation from the overall architecture, supporting heterogeneous task requirements. Optimal viewpoints are defined through systematic sampling and evaluation of candidate viewpoints, after which large-scale training datasets are constructed via domain randomization. Moreover, a multimodal optimal viewpoint prediction network is developed, leveraging cross-attention to align and fuse multimodal features and directly predict camera pose adjustments. The proposed framework is instantiated in robotic grasping under viewpoint-constrained environments. Experimental results demonstrate that active perception guided by the framework significantly improves grasp success rates. Notably, real-world evaluations achieve nearly double the grasp success rate and enable seamless sim-to-real transfer without additional fine-tuning, demonstrating the effectiveness of the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。