arXiv:2411.16804cs.CV2024-11被引 4

通过轨迹控制生成物体交互视频,提升真实感与可控性。

InTraGen: Trajectory-controlled Video Generation for Object Interactions

  • 引入多模态交互编码与物体ID注入机制,增强物体间互动建模。
  • 在4个新数据集上验证,生成视频在视觉质量和轨迹准确性上均提升。
  • 提供全新轨迹质量评估指标,适合研究视频生成与物理模拟的开发者。

视频生成技术的进步显著提升了生成场景的真实感与质量,推动了将视频生成用作世界模拟器的工具开发。文本到视频(T2V)生成允许仅通过文本描述创建视频,但受限于文本的模糊性和时间信息不足,研究者探索了轨迹引导等额外控制信号以提高生成精度。然而,当前缺乏有效评估T2V模型生成多物体交互能力的方法。为此,我们提出InTraGen,一种基于轨迹的物体交互场景生成框架。构建4个新数据集,并设计新型轨迹质量评估指标以衡量性能。为实现物体交互,提出多模态交互编码管道与物体ID注入机制,增强物体-环境间的动态交互。实验表明,该方法在视觉保真度和定量指标上均有提升。代码与数据集已公开于https://github.com/insait-institute/InTraGen。

原文摘要 · Abstract (English)

Advances in video generation have significantly improved the realism and quality of created scenes. This has fueled interest in developing intuitive tools that let users leverage video generation as world simulators. Text-to-video (T2V) generation is one such approach, enabling video creation from text descriptions only. Yet, due to the inherent ambiguity in texts and the limited temporal information offered by text prompts, researchers have explored additional control signals like trajectory-guided systems, for more accurate T2V generation. Nonetheless, methods to evaluate whether T2V models can generate realistic interactions between multiple objects are lacking. We introduce InTraGen, a pipeline for improved trajectory-based generation of object interaction scenarios. We propose 4 new datasets and a novel trajectory quality metric to evaluate the performance of the proposed InTraGen. To achieve object interaction, we introduce a multi-modal interaction encoding pipeline with an object ID injection mechanism that enriches object-environment interactions. Our results demonstrate improvements in both visual fidelity and quantitative performance. Code and datasets are available at https://github.com/insait-institute/InTraGen

视频生成轨迹控制物体交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。