arXiv:2510.03135cs.CVcs.RO2025-10AAAI被引 12

用轨迹预测生成人机交互视频,无需密集掩码标注。

Mask2IV: Interaction-Centric Video Generation via Mask Trajectories

  • 分两阶段:先预测动作与物体运动轨迹,再据此生成视频
  • 在两个新基准上实现更逼真、更可控的交互视频生成
  • 支持文本或位置提示控制,适合机器人学习与具身智能研究

生成以互动为核心的视频(如人类或机器人与物体交互)对具身智能至关重要,可为机器人学习、操作策略训练和可及性推理提供丰富的视觉先验。然而,现有方法难以建模此类复杂动态交互。尽管近期研究表明掩码可作为有效控制信号提升生成质量,但获取密集精确的掩码标注仍是实际应用的主要挑战。为此,我们提出Mask2IV,一种专为交互式视频生成设计的新框架。该框架采用解耦的两阶段流程:首先预测演员与物体的合理运动轨迹,然后基于这些轨迹生成视频。此设计无需用户输入密集掩码,同时保持对交互过程的灵活控制。此外,Mask2IV支持多样且直观的控制方式,用户可通过动作描述或空间位置线索指定目标物体并引导运动轨迹。为支持系统化训练与评估,我们构建了两个基准数据集,涵盖人类-物体交互与机器人操作场景中的多种动作与物体类别。大量实验表明,该方法在视觉真实性和可控性方面优于现有基线。

原文摘要 · Abstract (English)

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training, and affordance reasoning. However, existing methods often struggle to model such complex and dynamic interactions. While recent studies show that masks can serve as effective control signals and enhance generation quality, obtaining dense and precise mask annotations remains a major challenge for real-world use. To overcome this limitation, we introduce Mask2IV, a novel framework specifically designed for interaction-centric video generation. It adopts a decoupled two-stage pipeline that first predicts plausible motion trajectories for both actor and object, then generates a video conditioned on these trajectories. This design eliminates the need for dense mask inputs from users while preserving the flexibility to manipulate the interaction process. Furthermore, Mask2IV supports versatile and intuitive control, allowing users to specify the target object of interaction and guide the motion trajectory through action descriptions or spatial position cues. To support systematic training and evaluation, we curate two benchmarks covering diverse action and object categories across both human-object interaction and robotic manipulation scenarios. Extensive experiments demonstrate that our method achieves superior visual realism and controllability compared to existing baselines.

视频生成交互建模轨迹预测具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。