用3D空间感知提升机器人在不同光照下的视觉鲁棒性
Spatially Visual Perception for End-to-End Robotic Learning
- 结合图像增强与单目深度估计,构建视频空间感知框架
- 在多变光照下成功率达92%,旧模型性能大幅下降
- 适合做端到端机器人学习的低成本鲁棒视觉系统
模仿学习在机器人控制和具身智能中展现出巨大潜力,但应对多种安装摄像头观测的鲁棒泛化仍是关键挑战。本文提出一种基于视频的空间感知框架,利用3D空间表示应对环境变化,尤其关注光照变化。方法融合了新型图像增强技术AugBlender与在互联网规模数据上训练的先进单目深度估计模型。该系统显著提升了动态场景中的鲁棒性和适应性。实验表明,在多样化的相机曝光条件下,本方法的成功率显著优于先前模型,后者在复杂光照下出现性能崩溃。研究结果表明,基于视频的空间感知模型能有效推动端到端机器人学习的鲁棒性,为具身智能提供可扩展、低成本的解决方案。
原文摘要 · Abstract (English)
Recent advances in imitation learning have shown significant promise for robotic control and embodied intelligence. However, achieving robust generalization across diverse mounted camera observations remains a critical challenge. In this paper, we introduce a video-based spatial perception framework that leverages 3D spatial representations to address environmental variability, with a focus on handling lighting changes. Our approach integrates a novel image augmentation technique, AugBlender, with a state-of-the-art monocular depth estimation model trained on internet-scale data. Together, these components form a cohesive system designed to enhance robustness and adaptability in dynamic scenarios. Our results demonstrate that our approach significantly boosts the success rate across diverse camera exposures, where previous models experience performance collapse. Our findings highlight the potential of video-based spatial perception models in advancing robustness for end-to-end robotic learning, paving the way for scalable, low-cost solutions in embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。