用少量关键点控制人类与物体互动视频生成,更灵活真实。
DISPLAY: Directable Human-Object Interaction Video Generation via Sparse Motion Guidance and Multi-Task Auxiliary
- 仅需手腕坐标和物体框实现精准控制
- 在稀疏引导下仍保持高画质与物理一致性
- 适合需要自由交互的动画、虚拟人场景
以人为中心的视频生成技术快速发展,但现有方法难以生成可控且符合物理规律的人-物交互(HOI)视频。现有工作依赖密集控制信号、模板视频或精心设计的文本提示,限制了灵活性和对新物体的泛化能力。本文提出DISPLAY框架,仅使用手腕关节坐标和无形状依赖的物体边界框作为稀疏运动引导,缓解人与物表示不平衡问题,实现直观用户控制。为提升稀疏条件下的生成保真度,提出物体强化注意力机制以增强物体鲁棒性。针对高质量HOI数据稀缺问题,构建多任务辅助训练策略及专用数据清洗流程,使模型可同时受益于可靠HOI样本与辅助任务。大量实验表明,该方法在多样化任务中均能生成高保真、可控制的HOI视频。
原文摘要 · Abstract (English)
Human-centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human-Object Interaction (HOI) videos. Existing works rely on dense control signals, template videos, or carefully crafted text prompts, which limit flexibility and generalization to novel objects. We introduce a framework, namely DISPLAY, guided by Sparse Motion Guidance, composed only of wrist joint coordinates and a shape-agnostic object bounding box. This lightweight guidance alleviates the imbalance between human and object representations and enables intuitive user control. To enhance fidelity under such sparse conditions, we propose an Object-Stressed Attention mechanism that improves object robustness. To address the scarcity of high-quality HOI data, we further develop a Multi-Task Auxiliary Training strategy with a dedicated data curation pipeline, allowing the model to benefit from both reliable HOI samples and auxiliary tasks. Comprehensive experiments show that our method achieves high-fidelity, controllable HOI generation across diverse tasks. The project page can be found at \href{https://mumuwei.github.io/DISPLAY/}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。