用稀疏轨迹生成可控的人物-物体交互视频,效果更真实。
VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification
- 先将稀疏轨迹补全为稠密的交互掩码序列
- 在扩散模型中用稠密掩码控制生成,实现高保真交互
- 适合需要精准动作控制的动画与仿真场景
从稀疏轨迹生成逼真人-物交互视频极具挑战,因人体与物体的交互动态复杂且具实例特性。现有可控视频生成方法存在权衡:稀疏控制(如关键点轨迹)易指定但缺乏实例感知;稠密信号(如光流、深度或3D网格)信息丰富但获取成本高。本文提出VHOI,一个两阶段框架:首先将稀疏轨迹密集化为人-物交互掩码序列,再以这些稠密掩码微调视频扩散模型。引入一种新型人-物交互感知运动表示,通过颜色编码区分人体、物体及身体部位的动态。该设计将人体先验融入条件信号,增强模型对真实交互动态的理解与生成能力。实验表明,VHOI在可控人-物交互视频生成上达到当前最优性能。该方法不仅适用于仅交互场景,还能端到端生成完整人类导航至物体交互的过程。项目页面:https://vcai.mpi-inf.mpg.de/projects/vhoi/。
原文摘要 · Abstract (English)
Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the complexity. Existing controllable video generation approaches face a trade-off: sparse controls like keypoint trajectories are easy to specify but lack instance-awareness, while dense signals such as optical flow, depths or 3D meshes are informative but costly to obtain. We propose VHOI, a two-stage framework that first densifies sparse trajectories into HOI mask sequences, and then fine-tunes a video diffusion model conditioned on these dense masks. We introduce a novel HOI-aware motion representation that uses color encodings to distinguish not only human and object motion, but also body-part-specific dynamics. This design incorporates a human prior into the conditioning signal and strengthens the model's ability to understand and generate realistic HOI dynamics. Experiments demonstrate state-of-the-art results in controllable HOI video generation. VHOI is not limited to interaction-only scenarios and can also generate full human navigation leading up to object interactions in an end-to-end manner. Project page: https://vcai.mpi-inf.mpg.de/projects/vhoi/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。