轻量级网络实现2K实时立体视频转换,无需深度图输入
Lightweight Multiplane Images Network for Real-Time Stereoscopic Conversion from Planar Video
- 用轻量深度语义分支隐式提取深度特征
- 训练时用粗到精辅助分支,推理时仅保留轻量结构
- 相比最先进方法提速40倍,适合实时3D内容生成
随着无眼镜3D屏和虚拟现实设备的快速发展,立体视频资源匮乏问题凸显,立体转换成为关键任务。现有算法难以兼顾重建质量与推理效率。本文提出一种基于多平面图像(MPI)的平面视频实时立体转换网络,包含细节分支生成MPI和深度语义分支感知深度信息。不同于依赖显式深度图的方法,该方法通过轻量级深度语义分支隐式提取深度特征。为优化轻量分支,采用重训练、轻推理策略,设计仅在训练阶段使用的粗到精辅助分支。同时简化了立体转换场景下的MPI渲染流程,进一步加速推理。实验表明,该方法在2K分辨率下可实现接近顶尖模型的性能,并支持实时推理。相较于最先进方法TMPI,主观质量相当,但推理速度提升超过40倍。
原文摘要 · Abstract (English)
With the rapid development of stereoscopic display technologies, especially glasses-free 3D screens, and virtual reality devices, stereoscopic conversion has become an important task to address the lack of high-quality stereoscopic image and video resources. Current stereoscopic conversion algorithms typically struggle to balance reconstruction performance and inference efficiency. This paper proposes a planar video real-time stereoscopic conversion network based on multi-plane images (MPI), which consists of a detail branch for generating MPI and a depth-semantic branch for perceiving depth information. Unlike models that depend on explicit depth map inputs, the proposed method employs a lightweight depth-semantic branch to extract depth-aware features implicitly. To optimize the lightweight branch, a heavy training but light inference strategy is adopted, which involves designing a coarse-to-fine auxiliary branch that is only used during the training stage. In addition, the proposed method simplifies the MPI rendering process for stereoscopic conversion scenarios to further accelerate the inference. Experimental results demonstrate that the proposed method can achieve comparable performance to some state-of-the-art (SOTA) models and support real-time inference at 2K resolution. Compared to the SOTA TMPI algorithm, the proposed method obtains similar subjective quality while achieving over $40\times$ inference acceleration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。