arXiv:2606.13655cs.CVcs.GR2026-06

用单目或稀疏多视角视频生成动态4D人体,无需几何先验。

Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

论文配图:Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction
图 1 · 摘自论文原文
  • 通过相对相机位姿编码控制生成,不依赖骨骼或深度图。
  • 在DNA-Rendering和ActorsHQ数据集上优于现有方法,支持动物泛化。
  • 适合游戏、AR/VR等场景的低成本4D内容生成,可直接对接3D高斯溅射重建。

我们提出Flex4DHuman,一种多视角视频扩散模型,仅需相对相机位姿条件,即可将单目或稀疏多视角动态主体视频转换为同步的密集多视角视频。不同于依赖骨骼、深度图、法向或渲染目标视图几何的先前方法,Flex4DHuman无需显式几何先验,而是通过相对相机位姿的位置编码进行条件生成。生成视频可直接输入下游重建流程,构建动态4D高斯溅射。基于Wan 2.1 1.3B文生视频模型,保留主干结构,并通过五轴位置编码扩展时空RoPE,加入视图索引与连续SE(3)相对相机几何。采用三阶段课程训练:姿态跟随、灵活参考到目标视图生成、时间滚动。为支持时间滚动,训练中使用干净的历史目标视图标记;还添加多视角字幕以实现测试时文本控制。结合现成的4D高斯溅射阶段,该框架可将单目静态相机视频提升为动态4D高斯溅射。在DNA-Rendering和ActorsHQ上的实验表明,Flex4DHuman超越现有最先进方法,且经人类-动物混合训练后可泛化至动物类别。这些能力使该模型成为从日常单目视频实现可扩展4D内容创作的重要一步,适用于仿真、游戏、AR/VR及视频重拍。

原文摘要 · Abstract (English)

We present Flex4DHuman, a multi-view video diffusion model that transforms a monocular or sparse multi-view video of a dynamic subject into synchronized dense multi-view videos using only relative camera-pose conditioning. Unlike prior human-centric methods that rely on skeletons, depth maps, normals, or rendered target-view geometry, Flex4DHuman requires no explicit geometry priors and instead conditions generation through relative camera-pose positional encoding. The generated videos can be directly ingested by downstream reconstruction pipelines to create dynamic 4D Gaussian splats. Built on the Wan 2.1 1.3B text-to-video model, Flex4DHuman preserves the backbone architecture and encodes camera and view information through a five-axis positional encoding that extends spatio-temporal RoPE with view indices and continuous SE(3) relative camera geometry. A three-stage curriculum progressively trains the model for pose following, flexible reference-to-target view generation, and temporal rollout. To support temporal rollout, we train with clean historical target-view tokens. We also add multi-view captions to enable test-time text control. Combined with an off-the-shelf 4D Gaussian Splatting stage, our framework lifts monocular static-camera videos into dynamic 4D Gaussian splats. Experiments on DNA-Rendering and ActorsHQ show that Flex4DHuman surpasses prior state-of-the-art methods, while the same formulation generalizes to animal categories after mixed human-animal training. These capabilities make Flex4DHuman a practical step toward scalable 4D content creation from casual monocular videos for simulation, gaming, AR/VR, and video re-shooting.

4D重建视频生成扩散模型高斯溅射

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。