arXiv:2603.26757cs.RO2026-03被引 2

多视角演示提升机器人抓取性能,还能突破数据规模瓶颈。

Beyond Viewpoint Generalization: What Multi-View Demonstrations Offer and How to Synthesize Them for Robot Manipulation?

  • 通过多视角数据学习更鲁棒的视觉表征,优化动作决策分布。
  • 视角数量非单调提升,存在最优覆盖范围,能突破单视角性能天花板。
  • 提出RoboNVS合成新视角视频,适用于仿真与真实场景部署。

多视角示范是否真正提升机器人操作性能,还是仅增强跨视角鲁棒性?我们系统性地量化了多视角数据在机器人操作中的性能增益、扩展规律及内在机制。受控实验表明,在固定和随机背景条件下,多视角示范均显著提升单视角策略的成功率与泛化能力。性能随视角覆盖率变化呈非单调趋势,揭示有效区间而非简单的“越多越好”。值得注意的是,多视角数据打破了单视角数据集的扩展限制,在饱和后仍持续提升性能上限。机制分析显示,多视角学习促进与操作相关的视觉表征,更好对齐动作头与特征分布,并降低过拟合。鉴于多视角数据在大规模机器人数据集中稀缺且现实采集困难,我们提出RoboNVS——一种几何感知的自监督框架,可从单目输入生成新视角视频。生成数据在仿真与真实环境中均持续提升下游策略表现。

原文摘要 · Abstract (English)

Does multi-view demonstration truly improve robot manipulation, or merely enhance cross-view robustness? We present a systematic study quantifying the performance gains, scaling behavior, and underlying mechanisms of multi-view data for robot manipulation. Controlled experiments show that, under both fixed and randomized backgrounds, multi-view demonstrations consistently improve single-view policy success and generalization. Performance varies non-monotonically with view coverage, revealing effective regimes rather than a simple "more is better" trend. Notably, multi-view data breaks the scaling limitation of single-view datasets and continues to raise performance ceilings after saturation. Mechanistic analysis shows that multi-view learning promotes manipulation-relevant visual representations, better aligns the action head with the learned feature distribution, and reduces overfitting. Motivated by the importance of multi-view data and its scarcity in large-scale robotic datasets, as well as the difficulty of collecting additional viewpoints in real world settings, we propose RoboNVS, a geometry-aware self-supervised framework that synthesizes novel-view videos from monocular inputs. The generated data consistently improves downstream policies in both simulation and real-world environments.

机器人操作多视角学习数据合成自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。