让单视角人脸视频生成稳定可控的动态镜头,解决畸变问题。
FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning

- 用专为人脸设计的尺度感知相机表示,避免依赖3D先验
- 在Ava-256等数据集上实现高质量、身份一致的镜头控制
- 适合需要自然镜头运动的人像视频生成场景
我们提出FaceCam系统,可基于单视角人像视频生成自定义相机轨迹的视频。现有基于大模型的相机控制方法常因尺度模糊或3D重建误差,在人像视频中产生几何失真和视觉伪影。为此,我们设计了一种面向人脸的尺度感知相机变换表示,实现确定性条件生成,无需依赖3D先验。模型在多视角棚拍数据与真实环境单视角视频上联合训练,并引入合成相机运动与多帧拼接两种数据生成策略,使固定训练相机能泛化至推理时的动态连续轨迹。在Ava-256数据集及多样真实视频上的实验表明,FaceCam在相机可控性、视觉质量、身份与动作保持方面均表现优异。
原文摘要 · Abstract (English)
We introduce FaceCam, a system that generates video under customizable camera trajectories for monocular human portrait video input. Recent camera control approaches based on large video-generation models have shown promising progress but often exhibit geometric distortions and visual artifacts on portrait videos due to scale-ambiguous camera representations or 3D reconstruction errors. To overcome these limitations, we propose a face-tailored scale-aware representation for camera transformations that provides deterministic conditioning without relying on 3D priors. We train a video generation model on both multi-view studio captures and in-the-wild monocular videos, and introduce two camera-control data generation strategies: synthetic camera motion and multi-shot stitching, to exploit stationary training cameras while generalizing to dynamic, continuous camera trajectories at inference time. Experiments on Ava-256 dataset and diverse in-the-wild videos demonstrate that FaceCam achieves superior performance in camera controllability, visual quality, identity and motion preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。