仅用少数关键帧即可生成逼真人物物品交互视频,大幅降低标注成本。
SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation

- 通过稀疏关键帧与时间控制旋转位置编码实现精准时序定位。
- 利用多模态大模型提取运动先验,生成符合物理规律的中间帧。
- 构建5000条高质量稀疏控制数据集,适合电商视频生成场景。
人-物交互(HOI)视频生成旨在合成人类操作多样化物体的真实视频,是AI驱动直播电商的潜在方向。该领域主要挑战在于建模精细物理动态及人体手部与物体间的复杂时空协同。现有方法通常依赖密集时序引导(如逐帧手物姿态序列),严格控制交互过程,但导致标注成本高且影响运动多样性。为此,我们提出SparseCtrl-HOI,一种新型稀疏时序控制框架,仅需在指定时间戳的关键帧捕捉交互状态。具体而言,采用时间控制旋转位置编码(TiRoPE)机制对关键帧进行时序锚定,同时保持空间完整性。为调控中间帧的动态演化,提出运动先验注入模块,借助多模态大语言模型(MLLMs)提取高层运动先验,使模型能合理推断逻辑与物理上合理的过渡过程。此外,我们构建了名为SparseHOI-5K的高质量、丰富标注的数据集,支持稀疏时序控制的HOI视频生成。全面评估表明,该方法显著降低标注开销,同时生成更优的直播电商视频。代码与数据集已公开于https://mpi-lab.github.io/SparseCtrl-HOI。
原文摘要 · Abstract (English)
Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenue for AI-driven live streaming e-commerce. A primary obstacle in this domain lies in the complexity of modeling fine-grained physical dynamics and the intricate spatial-temporal coordination between human hands and objects. Existing approaches to this problem typically rely on dense temporal guidance, e.g., frame-wise hand-object pose sequences, to strictly control the interaction process. However, such dense guidance incurs high annotation costs and affects motion synthesis diversity. To overcome these limitations, we introduce SparseCtrl-HOI, a novel sparse temporal control framework for HOI video generation. It requires only a few keyframes that capture interaction states at designated timestamps. Specifically, we employ a Time-Controlled Rotary Positional Embedding (TiRoPE) mechanism to temporally anchor these keyframes while preserving their spatial integrity. Subsequently, to govern the dynamics across intermediate frames, we propose a Motion Prior Injection Module that leverages Multimodal Large Language Models (MLLMs) to extract high-level motion priors. This empowers the model to hallucinate logically and physically plausible transitions. Furthermore, we build SparseHOI-5K, a high-quality and richly annotated dataset for HOI video generation with sparse temporal control. Comprehensive evaluations confirm that our method substantially reduces annotation overhead while synthesizing superior live-streaming e-commerce videos. Both our code and dataset are publicly available at https://mpi-lab.github.io/SparseCtrl-HOI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。