让机器人通过主动感知与操作协同完成复杂任务
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
- 相机与操作动作解耦训练,分步提升能力
- 真实场景任务成功率最高提升31.25%
- 适合研究视觉语言动作模型的机器人开发者
主动感知与操作对机器人理解复杂场景至关重要。现有方法难以统一语义驱动的主动感知与鲁棒、视角不变的操作执行。我们提出SaPaVe,一种端到端框架,以数据高效方式联合学习这两种能力。方法将相机控制与操作动作解耦,而非共享动作空间,并采用自下而上的训练策略:先在大规模数据集上训练语义相机控制,再用混合数据联合优化两类动作。为支持该框架,我们构建了包含20万组图像-语言-相机运动对的ActiveViewPose-200K数据集,以及一个3D几何感知模块,提升动态视角下的执行鲁棒性。此外,我们提出了首个超越固定视角设置的主动操作评估基准ActiveManip-Bench。仿真与真实环境中的大量实验表明,SaPaVe优于GR00T N1和π₀等近期模型,在真实任务中成功率最高提升31.25%。结果表明,通过解耦但协调的策略训练紧密耦合的感知与执行,可实现高效且通用的主动操作。
原文摘要 · Abstract (English)
Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven active perception with robust, viewpoint-invariant execution. We propose SaPaVe, an end-to-end framework that jointly learns these capabilities in a data-efficient manner. Our approach decouples camera and manipulation actions rather than placing them in a shared action space, and follows a bottom-up training strategy: we first train semantic camera control on a large-scale dataset, then jointly optimize both action types using hybrid data. To support this framework, we introduce ActiveViewPose-200K, a dataset of 200k image-language-camera movement pairs for semantic camera movement learning, and a 3D geometry-aware module that improves execution robustness under dynamic viewpoints. We also present ActiveManip-Bench, the first benchmark for evaluating active manipulation beyond fixed-view settings. Extensive experiments in both simulation and real-world environments show that SaPaVe outperforms recent vision-language-action models such as GR00T N1 and \(π_0\), achieving up to 31.25\% higher success rates in real-world tasks. These results show that tightly coupled perception and execution, when trained with decoupled yet coordinated strategies, enable efficient and generalizable active manipulation. Project page: https://lmzpai.github.io/SaPaVe
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。