单图实时生成人像正面视图,兼顾真实感与结构准确
Real-Time Human Frontal View Synthesis from a Single Image
- 通过级联学习直接建模粗粒度几何特征,再精修纹理
- 在24帧/秒下实现实时推理,显著优于现有方法
- 无需外部模型,适合沉浸式3D远程通信场景
从单张图像实现逼真的人体新视角合成对普及沉浸式3D远程通信至关重要,可避免复杂的多相机部署。然而,当前以渲染为核心的 方法虽注重视觉保真度,却缺乏显式的几何理解,在面部和手部等复杂区域易产生时序不稳定性。而以人为中心的框架因依赖辅助模型提供结构先验,常面临内存瓶颈,难以实现实时性能。为此,我们提出PrismMirror,一种面向单图实时正面视图合成的几何引导框架。该模型摒弃外部几何建模,聚焦正面视图生成,优化远程通信的视觉完整性。具体地,PrismMirror引入新颖的级联学习策略,实现从粗到细的几何特征学习:先直接学习粗粒度几何特征(如SMPL-X网格与点云),再通过渲染监督精修纹理。为实现实时效率,我们将该统一框架压缩为轻量级线性注意力模型。值得注意的是,PrismMirror是首个在单目条件下实现24帧/秒实时推理的人体正面视图合成模型,在视觉真实性和结构准确性上均显著优于此前方法。
原文摘要 · Abstract (English)
Photorealistic human novel view synthesis from a single image is crucial for democratizing immersive 3D telepresence, eliminating the need for complex multi-camera setups. However, current rendering-centric methods prioritize visual fidelity over explicit geometric understanding and struggle with intricate regions like faces and hands, leading to temporal instability. Meanwhile, human-centric frameworks suffer from memory bottlenecks since they typically rely on an auxiliary model to provide informative structural priors for geometric modeling, which limits real-time performance. To address these challenges, we propose PrismMirror, a geometry-guided framework for instant frontal view synthesis from a single image. By avoiding external geometric modeling and focusing on frontal view synthesis, our model optimizes visual integrity for telepresence. Specifically, PrismMirror introduces a novel cascade learning strategy that enables coarse-to-fine geometric feature learning. It first directly learns coarse geometric features, such as SMPL-X meshes and point clouds, and then refines textures through rendering supervision. To achieve real-time efficiency, we distill this unified framework into a lightweight linear attention model. Notably, PrismMirror is the first monocular human frontal view synthesis model that achieves real-time inference at 24 FPS, significantly outperforming previous methods in both visual authenticity and structural accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。