通过智能选位摄像头,让机械臂在不同视角下都能稳定操作。
Viewpoint-Agnostic Manipulation Policies with Strategic Vantage Selection
- 用信息增益优化动态选摄像头位置,避免盲目试错。
- 仅需少量微调就使成功率提升25%,实测表现更稳。
- 适合想提升视觉抓取鲁棒性的机器人研究者和工程师。
由于基于视觉的操控策略通常从单一视角数据中训练,部署时视角变化会导致性能下降。直接聚合大量随机视角的示范不仅成本高,还可能因视觉差异过大而干扰学习。我们提出Vantage框架,通过在少量精心选择的相机位姿上微调预训练策略,实现对视角不变的行为。不同于耗时的暴力搜索,Vantage将相机布局建模为连续空间中的信息增益优化问题,平衡了探索新视角与利用已有有效视角,同时提供收敛性和鲁棒性的理论保证。在多种任务和策略类型中,Vantage均显著优于固定、网格或随机的数据选择策略,仅需少量微调步骤即可提升视角变换下的成功率。模拟与真实场景实验表明,该方法使扩散策略的任务成功率提升25%,并在动态摄像头设置下表现出稳健增益。
原文摘要 · Abstract (English)
Since vision-based manipulation policies are typically trained from data gathered from a single viewpoint, their performance drops when the view changes during deployment. Naively aggregating demonstrations from numerous random views is not only costly but also known to destabilize learning, as excessive visual diversity acts as noise. We present Vantage, a viewpoint selection framework to fine-tune any pre-trained policy on a small, strategically set of camera poses to induce viewpoint-agnostic behavior. Instead of relying on costly brute-force search over viewpoints, Vantage formulates camera placement as an information gain optimization problem in a continuous space. This approach balances exploration of novel poses with exploitation of promising ones, while also providing theoretical guarantees about convergence and robustness. Across manipulation tasks and policy families, Vantage consistently improves success under viewpoint shifts compared to fixed, grid, or random data selection strategies with only a handful of fine-tuning steps. Experiments conducted on simulated and real-world setups show that Vantage increases the task success rate by 25% for diffusion policies, and yields robust gains in dynamic-camera settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。