arXiv:2512.13392cs.CV2025-12

对比了视频生成中运动与遮挡控制的四种设计组合,发现跟踪+图像引导效果最佳。

A Study of the Design Space of Motion and Disocclusion Control in Video Generation

  • 分运动与遮挡控制方式,构建合成与真实场景双类基准数据集
  • 追踪引导提升运动准确性,图像引导增强新暴露内容保真度
  • 动态新显物体的外观与运动联合控制仍是难点,适合视频生成研究者参考

当物体移动导致原本隐藏的内容显露时,即发生遮挡恢复(disocclusion)。现有图像到视频方法在运动与遮挡控制上采用不同形式,但其算法设计空间尚未被充分理解。本文系统探索了运动指定(基于文本或基于追踪)与遮挡恢复指定(基于文本或基于图像)构成的设计空间。为支持该研究,我们构建了一个包含合成场景与真实场景的基准数据集,涵盖刚性与可变形运动,以及静态与动态的新显物体。我们在该设计空间上比较现有及适配方法,分析不同控制策略在处理运动中遮挡恢复时的优劣。研究发现:基于追踪的引导能提升运动准确性,基于图像的引导则改善新暴露内容的保真度。然而,在新显物体自身后续运动的场景下,同时控制其外观与动态仍是一个具有挑战性的开放问题。

原文摘要 · Abstract (English)

Disocclusion occurs when object movement reveals previously hidden content. Existing image-to-video methods provide different forms of motion and disocclusion control, yet the design space they span remains poorly understood. In this work, we systematically explore an algorithmic design space spanning motion specification (text-based versus tracking-based) and disocclusion specification (text-based versus image-based). To facilitate this study, we construct a benchmark consisting of both synthetic scenes and captured scenes, covering rigid and deformable motions, as well as static and dynamic disoccluded objects. We compare existing and adapted methods across this design space to analyze the strengths and limitations of different control strategies for handling disocclusions during motion. Our study reveals that tracking-based guidance improves motion accuracy, while image-based guidance improves the fidelity of newly revealed content. However, jointly controlling the appearance and dynamics of disoccluded objects remains a challenging open problem, particularly in scenarios where newly revealed objects undergo their own motion after becoming visible.

视频生成运动控制遮挡恢复生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。