研究鱼眼相机在机器人抓取中的表现,发现视野优势需环境复杂度支撑。
Rethinking Camera Choice: An Empirical Study on Fisheye Camera Properties in Robotic Manipulation
- 通过实证分析鱼眼相机在模仿学习中的性能,验证其广角优势
- 环境多样性足够时,鱼眼训练策略泛化能力更强,可提升场景适应性
- 提出随机缩放增强法,解决跨设备迁移中的尺度过拟合问题
鱼眼相机因超广视角在机器人抓取中应用日益广泛,但其对策略学习的下游影响尚未系统理解。本文首次开展全面实证研究,针对腕装鱼眼相机在模仿学习中的特性进行分析。在仿真与真实世界中,围绕空间定位、场景泛化与硬件泛化三个核心问题展开实验。结果表明:(1) 超广视角显著提升空间定位能力,但该优势依赖环境视觉复杂度;(2) 在充分环境多样性的训练下,鱼眼模型虽在简单场景易过拟合,却展现出更优的场景泛化能力;(3) 未经处理的跨相机迁移会失败,根源为尺度过拟合,采用简单的随机缩放增强(RSA)策略可有效改善硬件泛化性能。研究成果为大规模鱼眼数据集的构建与高效利用提供明确指导。更多结果与视频见 https://robo-fisheye.github.io/
原文摘要 · Abstract (English)
The adoption of fisheye cameras in robotic manipulation, driven by their exceptionally wide Field of View (FoV), is rapidly outpacing a systematic understanding of their downstream effects on policy learning. This paper presents the first comprehensive empirical study to bridge this gap, rigorously analyzing the properties of wrist-mounted fisheye cameras for imitation learning. Through extensive experiments in both simulation and the real world, we investigate three critical research questions: spatial localization, scene generalization, and hardware generalization. Our investigation reveals that: (1) The wide FoV significantly enhances spatial localization, but this benefit is critically contingent on the visual complexity of the environment. (2) Fisheye-trained policies, while prone to overfitting in simple scenes, unlock superior scene generalization when trained with sufficient environmental diversity. (3) While naive cross-camera transfer leads to failures, we identify the root cause as scale overfitting and demonstrate that hardware generalization performance can be improved with a simple Random Scale Augmentation (RSA) strategy. Collectively, our findings provide concrete, actionable guidance for the large-scale collection and effective use of fisheye datasets in robotic learning. More results and videos are available on https://robo-fisheye.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。