arXiv:2506.19840cs.CV2025-06被引 6

无需训练即可生成长视频中人与场景的自然互动,保持角色身份一致。

GenHSI: Controllable Generation of Human-Scene Interaction Videos

  • 分三阶段生成:写剧本、预可视化、动画,用2D扩散模型生成3D关键帧。
  • 通过接触提示和视觉语言模型推理,实现无多视角拟合的3D动作合成。
  • 仅需一张场景图和角色描述,就能生成连贯且符合物理交互的长视频。

大规模预训练视频扩散模型在视频生成方面展现出强大能力,但在生成包含丰富人-场景交互(HSI)的长视频时仍面临动态不真实、功能合理性不足、角色身份难以保持及高昂训练成本等问题。为此,我们提出GenHSI,一种无需训练的可控方法,用于生成具有3D感知能力的长时人-场景交互视频。受电影动画启发,我们将视频生成分为三个阶段:(1) 脚本编写,(2) 预可视化,(3) 动画生成。给定场景图像和角色描述,通过将复杂链式交互文本转化为原子动作,用于预可视化阶段生成3D关键帧。为合成合理的3D人体交互姿态,我们利用预训练的2D修复扩散模型,结合视图规范化生成合理2D交互,并通过接触线索和视觉语言模型(VLM)推理进行鲁棒迭代优化,实现3D扩展。基于这些3D关键帧,预训练视频扩散模型可生成具有一致性与合理动力学特性的长视频。我们首次在仅依赖场景与角色图像参考的情况下,实现无需训练的长序列链式人-场景交互视频生成。实验表明,该方法能有效保留场景内容与角色身份,生成符合物理交互逻辑的高质量视频。

原文摘要 · Abstract (English)

Large-scale pre-trained video diffusion models have exhibited remarkable capabilities in diverse video generation. However, existing solutions face several challenges in generating long videos with rich human-scene interactions (HSI), including unrealistic dynamics and affordance, lack of subject identity preservation, and the need for expensive training. To this end, we propose GenHSI, a training-free method for controllable generation of long HSI videos with 3D awareness. Taking inspiration from movie animation, we subdivide the video synthesis into three stages: (1) script writing, (2) pre-visualization, and (3) animation. Given an image of a scene and a character with a user description, we use these three stages to generate long videos that preserve human identity and provide rich and plausible HSI. Script writing converts a complex text prompt involving a chain of HSI into simple atomic actions that are used in the pre-visualization stage to generate 3D keyframes. To synthesize plausible human interaction poses in 3D keyframes, we utilize pre-trained 2D inpainting diffusion models to generate plausible 2D human interactions based on view canonicalization, which eliminates the need for multi-view fitting in previous works. We then extend these interactions to 3D using robust iterative optimization, informed by contact cues and reasoning from VLMs. Prompted by these 3D keyframes, the pretrained video diffusion models can better generate consistent long videos with plausible dynamics and affordance in a 3D-aware manner. We are the first to synthesize a long video sequence with a chain of HSI actions without training based on the image references of the scene and character. Experiments demonstrate that our method can generate HSI videos that effectively preserve scene content and character identity with plausible human-scene interaction from a single image scene.

视频生成人-场景交互扩散模型3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。