arXiv:2510.03550cs.CV2025-10被引 7

让用户随时拖动视频任意元素,实现流畅精准的交互式编辑。

Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!

  • 通过自适应修正潜空间分布,缓解拖拽导致的模型失准问题。
  • 在不需训练的情况下,实现毫秒级响应的实时拖拽效果。
  • 适合需要快速编辑视频内容的创作者与设计师使用。

实现对自回归视频扩散模型输出的流式、细粒度控制仍具挑战性,难以确保结果始终符合用户预期。为此,我们提出新任务stReaming drag-oriEnted interactiVe vidEo manipuLation(REVEL),使用户可随时对任意物体进行精细拖拽操作。与DragVideo和SG-I2V不同,REVEL统一了拖拽式视频编辑与动画生成,支持用户指定的平移、形变和旋转等效果。我们发现:拖拽引起的潜空间扰动会累积,导致严重分布漂移并中断操作;流式拖拽易受上下文帧干扰,产生视觉不自然结果。为此,我们提出无需训练的DragStream方法:其一,采用自适应分布自校正策略,利用邻近帧统计信息约束潜变量漂移;其二,设计空间-频率选择性优化机制,在充分利用上下文信息的同时,通过选择性传播视觉线索减轻干扰。该方法可无缝集成至现有自回归视频扩散模型中,大量实验验证了其有效性。

原文摘要 · Abstract (English)

Achieving streaming, fine-grained control over the outputs of autoregressive video diffusion models remains challenging, making it difficult to ensure that they consistently align with user expectations. To bridge this gap, we propose \textbf{stReaming drag-oriEnted interactiVe vidEo manipuLation (REVEL)}, a new task that enables users to modify generated videos \emph{anytime} on \emph{anything} via fine-grained, interactive drag. Beyond DragVideo and SG-I2V, REVEL unifies drag-style video manipulation as editing and animating video frames with both supporting user-specified translation, deformation, and rotation effects, making drag operations versatile. In resolving REVEL, we observe: \emph{i}) drag-induced perturbations accumulate in latent space, causing severe latent distribution drift that halts the drag process; \emph{ii}) streaming drag is easily disturbed by context frames, thereby yielding visually unnatural outcomes. We thus propose a training-free approach, \textbf{DragStream}, comprising: \emph{i}) an adaptive distribution self-rectification strategy that leverages neighboring frames' statistics to effectively constrain the drift of latent embeddings; \emph{ii}) a spatial-frequency selective optimization mechanism, allowing the model to fully exploit contextual information while mitigating its interference via selectively propagating visual cues along generation. Our method can be seamlessly integrated into existing autoregressive video diffusion models, and extensive experiments firmly demonstrate the effectiveness of our DragStream.

视频编辑交互式扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。