arXiv:2505.21491cs.CV2025-05NeurIPS被引 8

让图像中物体按轨迹进出画面,实现自然可控的视频生成。

Frame In-N-Out: Unbounded Controllable Image-to-Video Generation

论文配图:Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
图 1 · 摘自论文原文
  • 基于用户指定轨迹控制物体进出场景,保持身份一致性。
  • 新架构在可控性与画面连贯性上显著优于现有方法。
  • 适用于影视特效、创意设计等需要精准动作控制的场景。

可控性、时间连贯性和细节合成仍是视频生成的核心挑战。本文聚焦一种常见但研究不足的电影技巧——‘帧内进入’与‘帧外退出’。从图生视频出发,用户可控制图像中的物体按指定运动轨迹自然离开场景,或引入新身份进入场景。为此,我们构建了半自动标注的新数据集,提出高效的身份保持型运动可控视频扩散变换器架构,并设计了针对该任务的综合评估协议。实验表明,所提方法显著优于现有基线。

原文摘要 · Abstract (English)

Controllability, temporal coherence, and detail synthesis remain the most critical challenges in video generation. In this paper, we focus on a commonly used yet underexplored cinematic technique known as Frame In and Frame Out. Specifically, starting from image-to-video generation, users can control the objects in the image to naturally leave the scene or provide breaking new identity references to enter the scene, guided by a user-specified motion trajectory. To support this task, we introduce a new dataset that is curated semi-automatically, an efficient identity-preserving motion-controllable video Diffusion Transformer architecture, and a comprehensive evaluation protocol targeting this task. Our evaluation shows that our proposed approach significantly outperforms existing baselines.

视频生成可控生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。