系统梳理全景视觉的表示学习与优化方法,助力自动驾驶与VR发展
A Survey of Representation Learning, Optimization Strategies, and Applications for Omnidirectional Vision
- 从投影原理到表示学习,构建全景视觉深度学习框架
- 涵盖图像增强、3D几何估计等任务的系统性分类与分析
- 适合从事视觉算法、VR/AR开发的研究者参考
全景图像(ODI)具有360×180°视场,远超针孔相机,能捕捉更丰富的周围环境细节。近年来,消费级360相机普及与深度学习(DL)进展推动了全景视觉研究热潮。本文系统综述了深度学习在全景视觉中的最新进展,剖析其相较于传统透视图像的独特挑战。内容包括:(i) 全景成像原理与常见投影方式;(ii) 针对ODI的多样化表示学习方法;(iii) 专用于全景视觉的优化策略;(iv) 从图像增强(如生成、超分辨率)到3D几何与运动估计(如深度、光流估计)的任务分类体系;(v) 前沿应用(如自动驾驶、虚拟现实)及当前挑战与开放问题的讨论,旨在激发领域持续研究。
原文摘要 · Abstract (English)
Omnidirectional image (ODI) data is captured with a field-of-view of 360x180, which is much wider than the pinhole cameras and captures richer surrounding environment details than the conventional perspective images. In recent years, the availability of customer-level 360 cameras has made omnidirectional vision more popular, and the advance of deep learning (DL) has significantly sparked its research and applications. This paper presents a systematic and comprehensive review and analysis of the recent progress of DL for omnidirectional vision. It delineates the distinct challenges and complexities encountered in applying DL to omnidirectional images as opposed to traditional perspective imagery. Our work covers four main contents: (i) A thorough introduction to the principles of omnidirectional imaging and commonly explored projections of ODI; (ii) A methodical review of varied representation learning approaches tailored for ODI; (iii) An in-depth investigation of optimization strategies specific to omnidirectional vision; (iv) A structural and hierarchical taxonomy of the DL methods for the representative omnidirectional vision tasks, from visual enhancement (e.g., image generation and super-resolution) to 3D geometry and motion estimation (e.g., depth and optical flow estimation), alongside the discussions on emergent research directions; (v) An overview of cutting-edge applications (e.g., autonomous driving and virtual reality), coupled with a critical discussion on prevailing challenges and open questions, to trigger more research in the community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。