arXiv:2507.00334cs.CV2025-07

让视频模型学会看懂场景能做什么,自动生成符合环境行为的人像视频。

Populate-A-Scene: Affordance-Aware Human Video Generation

  • 从单张场景图推断人与环境的交互位置和动作,无需标注框或姿态
  • 生成视频中人物行为自然、外观协调且符合场景功能
  • 通过注意力热图揭示预训练模型隐含的环境可操作性理解能力

文本到视频模型能否被用作交互式世界模拟器?我们探索了文本到视频模型在感知环境可操作性方面的潜力,通过训练模型根据场景图像和人类动作描述,生成插入人物的视频,并确保行为一致、外观协调及场景可操作性。不同于以往工作,本方法仅依赖单张场景图,无需边界框或人体姿态等显式条件,即可推断出人物应插入的位置及其合理行为。对交叉注意力热图的深入分析表明,无需标注的可操作性数据集,也能从预训练视频模型中挖掘出其内在的环境可操作性感知能力。

原文摘要 · Abstract (English)

Can a video generation model be repurposed as an interactive world simulator? We explore the affordance perception potential of text-to-video models by teaching them to predict human-environment interaction. Given a scene image and a prompt describing human actions, we fine-tune the model to insert a person into the scene, while ensuring coherent behavior, appearance, harmonization, and scene affordance. Unlike prior work, we infer human affordance for video generation (i.e., where to insert a person and how they should behave) from a single scene image, without explicit conditions like bounding boxes or body poses. An in-depth study of cross-attention heatmaps demonstrates that we can uncover the inherent affordance perception of a pre-trained video model without labeled affordance datasets.

视频生成场景理解可操作性感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。