用4个90度间隔的稀疏视角视频,实现2K分辨率全身360°实时重建。
HiReFF: High-Resolution Feedforward Human Reconstruction from Uncalibrated Sparse-View Video

- 分两步:先解耦相机尺度歧义,再用高斯掩码清理前景。
- 2K渲染仅需0.5K主干计算,合成效率提升显著。
- 适合全向高分辨率人形视频流,无需标定相机。
无标定体积视频流对全息通信和AR/VR至关重要,但受限于稀疏视角输入下的时序一致性与计算效率。现有方法依赖每场景优化或标定相机,而近期前馈模型仅限于0.5K分辨率单帧生成。本文提出HiReFF,可从四个相隔90°的稀疏视角视频中,实现2K分辨率360°人体视频的前馈重建。框架将问题分解为两个关键任务:从稀疏视角视频中进行前景3D高斯重建,以及高效高分辨率合成。为实现前者,提出尺度同步相机标定以解决多视角监督中的尺度歧义,并设计高斯级前景掩码通过调节高斯参数重建干净前景。为实现高效高分辨率合成,采用高分辨率侧调优,在保持主干网络为0.5K的同时,通过补充特征增强高斯头,实现2K渲染,大幅降低计算开销。实验表明,HiReFF在高分辨率体积视频重建上显著优于现有方法。
原文摘要 · Abstract (English)
Uncalibrated volumetric video streaming for human reconstruction is essential for holographic communication and AR/VR, yet remains challenging due to the need for temporal consistency and computational efficiency from sparse-view inputs. Existing methods rely on per-scene optimization or calibrated cameras, while recent feed-forward models are limited to low-resolution (0.5K) single-frame synthesis. We present HiReFF, a feed-forward method for 2K-resolution 360° human video reconstruction from uncalibrated sparse-view videos. Our framework decomposes the problem into two key tasks: foreground 3D Gaussian reconstruction from sparse-view videos (four views separated by 90°) and computationally efficient high-resolution synthesis. To enable the former, we propose Scale-synchronized Camera Calibration to resolve scale ambiguity for multi-view supervision, and Gaussian-wise Foreground Masking to reconstruct clean foregrounds by modulating Gaussian parameters. For efficient high-resolution synthesis, our High-resolution Side-tuning achieves 2K rendering by augmenting the Gaussian head with supplementary features while keeping the backbone at 0.5K, drastically reducing computational overhead. Experiments demonstrate that HiReFF significantly outperforms existing methods in high-resolution streaming volumetric video reconstruction. https://iridescentjiang.github.io/HiReFF
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。