arXiv:2505.15800cs.CV2025-05被引 9

用新注意力机制提升数字人视频生成质量与一致性

Interspatial Attention for Efficient 4D Human Video Generation

  • 引入跨空间注意力机制,基于相对位置编码优化人体视频生成
  • 在大规模视频数据上训练,实现高保真、姿态与身份一致的4D视频合成
  • 适合需要精确控制相机和身体姿态的虚拟人应用开发

以可控方式生成逼真的数字人视频对众多应用至关重要。现有方法或依赖模板化3D表示,或基于新兴视频生成模型,但在生成单个或多个数字人时存在质量差、运动不连贯及身份失真问题。本文提出一种新型跨空间注意力(ISA)机制,作为现代扩散变换器(DiT)模型的可扩展组件。ISA采用专为人体视频生成设计的相对位置编码,结合自研视频变分自编码器,在大规模视频数据上训练潜空间ISA扩散模型。该模型在4D人体视频合成任务中达到当前最优表现,展现出卓越的运动连贯性与身份保持能力,并支持对相机视角和身体姿态的精准控制。代码与模型已公开:https://dsaurus.github.io/isa4d/

原文摘要 · Abstract (English)

Generating photorealistic videos of digital humans in a controllable manner is crucial for a plethora of applications. Existing approaches either build on methods that employ template-based 3D representations or emerging video generation models but suffer from poor quality or limited consistency and identity preservation when generating individual or multiple digital humans. In this paper, we introduce a new interspatial attention (ISA) mechanism as a scalable building block for modern diffusion transformer (DiT)--based video generation models. ISA is a new type of cross attention that uses relative positional encodings tailored for the generation of human videos. Leveraging a custom-developed video variation autoencoder, we train a latent ISA-based diffusion model on a large corpus of video data. Our model achieves state-of-the-art performance for 4D human video synthesis, demonstrating remarkable motion consistency and identity preservation while providing precise control of the camera and body poses. Our code and model are publicly released at https://dsaurus.github.io/isa4d/.

视频生成扩散模型人体建模注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。