用无监督方法修复广角视频人脸变形,兼顾结构与细节。
Beyond Wide-Angle Images: Structure-to-Detail Video Portrait Correction via Unsupervised Spatiotemporal Adaptation
- 结合Transformer与扩散模型,实现全局结构与局部细节的协同修正。
- 在无标签视频上通过时空一致性约束,保持面部修正的稳定与平滑。
- 适用于多人群、复杂光照场景,适合视频创作与实时处理需求。
广角相机虽广泛用于内容创作,但镜头边缘易产生人脸拉伸畸变,影响视觉效果。为此,我们提出名为ImagePC的结构到细节人脸修复模型,融合Transformer的长程感知能力与扩散模型的多步去噪机制,实现全局结构鲁棒性与局部细节精细化。针对视频标注成本高的问题,进一步将ImagePC拓展为无监督视频修复框架VideoPC,通过引入空间一致性与时间平滑性约束进行时空扩散自适应:空间上逼近期望畸变分布的伪标签,时间上利用反向光流生成矫正轨迹并平滑。相比ImagePC,VideoPC在无标签场景下仍能保持空间高质量修正,并有效抑制时序抖动。最后,构建了一个包含多人数、多光照、多样背景的大规模视频人像数据集,用于训练与评估。实验表明,该方法在定量与定性指标上均优于现有方案,显著提升广角视频中人脸的保真度与自然性。代码与数据集将公开。
原文摘要 · Abstract (English)
Wide-angle cameras, despite their popularity for content creation, suffer from distortion-induced facial stretching-especially at the edge of the lens-which degrades visual appeal. To address this issue, we propose a structure-to-detail portrait correction model named ImagePC. It integrates the long-range awareness of the transformer and multi-step denoising of diffusion models into a unified framework, achieving global structural robustness and local detail refinement. Besides, considering the high cost of obtaining video labels, we then repurpose ImagePC for unlabeled wide-angle videos (termed VideoPC), by spatiotemporal diffusion adaption with spatial consistency and temporal smoothness constraints. For the former, we encourage the denoised image to approximate pseudo labels following the wide-angle distortion distribution pattern, while for the latter, we derive rectification trajectories with backward optical flows and smooth them. Compared with ImagePC, VideoPC maintains high-quality facial corrections in space and mitigates the potential temporal shakes sequentially in blind scenarios. Finally, to establish an evaluation benchmark and train the framework, we establish a video portrait dataset with a large diversity in the number of people, lighting conditions, and background. Experiments demonstrate that the proposed methods outperform existing solutions quantitatively and qualitatively, contributing to high-fidelity wide-angle videos with stable and natural portraits. The codes and dataset will be available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。