轻量级单视图新视角生成,手机上实时运行且效果领先。
CheapNVS: Real-Time On-Device Narrow-Baseline Novel View Synthesis

- 用轻量模块替代复杂3D变形,结合相机位姿条件优化
- 在开放图像子集训练后,速度提升10倍、内存降低6%
- 适合移动端实时应用,三星平板实测超30帧/秒
单视图新视角生成因病态性问题而困难重重,传统方法常需庞大计算资源。本文提出 CheapNVS:一种基于新型高效多编码器/解码器设计的端到端窄基线单视图 NVS 方法,采用多阶段训练策略。CheapNVS 首先以轻量可学习模块近似复杂的 3D 图像扭曲,该模块受目标视角相机位姿嵌入条件控制;随后并行对遮挡区域进行修补,显著提升性能。在 Open Images 数据集子集上训练后,CheapNVS 在保持 6% 更低内存占用的同时,实现比现有最优方法更快 10 倍的速度,并可在移动设备上流畅运行,实测在 Samsung Tab 9+ 上达到超过 30 FPS。
原文摘要 · Abstract (English)
Single-view novel view synthesis (NVS) is a notorious problem due to its ill-posed nature, and often requires large, computationally expensive approaches to produce tangible results. In this paper, we propose CheapNVS: a fully end-to-end approach for narrow baseline single-view NVS based on a novel, efficient multiple encoder/decoder design trained in a multi-stage fashion. CheapNVS first approximates the laborious 3D image warping with lightweight learnable modules that are conditioned on the camera pose embeddings of the target view, and then performs inpainting on the occluded regions in parallel to achieve significant performance gains. Once trained on a subset of Open Images dataset, CheapNVS outperforms the state-of-the-art despite being 10 times faster and consuming 6% less memory. Furthermore, CheapNVS runs comfortably in real-time on mobile devices, reaching over 30 FPS on a Samsung Tab 9+.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。