用多视角表演捕捉定制视频扩散模型,实现角色一致与3D相机控制。
Virtually Being: Customizing Camera-Controllable Video Diffusion Models with Multi-View Performance Captures
- 通过4D高斯点云重渲染多视角表演数据,构建定制化训练集。
- 支持多主体生成、真实场景定制及运动布局控制,提升个性化精度。
- 适合虚拟制作团队,可高效组合已定制模型进行实时创作。
我们提出一种框架,通过新颖的定制化数据流水线,实现视频扩散模型中的多视角角色一致性与3D相机控制。利用体积捕捉表演数据,结合4D高斯点云(4DGS)和视频重光照模型生成多样相机轨迹与光照变化,对开源先进视频扩散模型进行微调。该方法在保持强多视角身份一致性的同时,实现了精确相机控制与光照适应性。框架还支持两种多主体生成方式:联合训练与噪声融合,后者可在推理时高效组合独立定制模型;同时支持场景与真实视频定制,以及对运动和空间布局的定制控制。大量实验表明,视频质量提升,个性化准确率更高,相机控制与光照适应能力显著增强,推动视频生成在虚拟制作中的应用。项目主页:https://eyeline-labs.github.io/Virtually-Being。
原文摘要 · Abstract (English)
We introduce a framework that enables both multi-view character consistency and 3D camera control in video diffusion models through a novel customization data pipeline. We train the character consistency component with recorded volumetric capture performances re-rendered with diverse camera trajectories via 4D Gaussian Splatting (4DGS), lighting variability obtained with a video relighting model. We fine-tune state-of-the-art open-source video diffusion models on this data to provide strong multi-view identity preservation, precise camera control, and lighting adaptability. Our framework also supports core capabilities for virtual production, including multi-subject generation using two approaches: joint training and noise blending, the latter enabling efficient composition of independently customized models at inference time; it also achieves scene and real-life video customization as well as control over motion and spatial layout during customization. Extensive experiments show improved video quality, higher personalization accuracy, and enhanced camera control and lighting adaptability, advancing the integration of video generation into virtual production. Our project page is available at: https://eyeline-labs.github.io/Virtually-Being.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。