arXiv:2511.14100cs.CV2025-11被引 3

让视频编辑理解模糊指令,通过推理找到要改的地方

Text-Driven Reasoning Video Editing via Reinforcement Learning on Digital Twin Representations

  • 用数字孪生保留视频时空关系,分离推理与生成
  • 基于强化学习训练,准确率比基线高12.3%
  • 适合需要语义级编辑的创作者和研究者

文本驱动视频编辑允许用户仅通过文本查询修改视频内容。现有方法需明确描述编辑目标的空间位置和时间范围,但当用户使用隐含查询(如语义属性或对象关系)时,这些要求变得不切实际。本文提出推理视频编辑任务,即模型需通过多跳推理理解隐含查询以推断编辑目标,首次提出解决该任务的RIVER模型。RIVER通过视频内容的数字孪生表示解耦推理与生成,该表示保留空间关系、时间轨迹和语义属性。大型语言模型联合处理此表示与隐含查询,执行多跳推理以确定修改方案,并输出结构化指令,指导基于扩散的编辑器完成像素级修改。RIVER训练采用强化学习,奖励机制评估推理准确性和生成质量。此外,本文构建了包含100个视频、519个隐含查询的RVEBenchmark,涵盖三个层级和类别的推理复杂度。RIVER在该基准上表现最佳,并在VegGIE和FiVE两个额外视频编辑基准上超越六种基线方法,达到当前最优水平。

原文摘要 · Abstract (English)

Text-driven video editing enables users to modify video content only using text queries. While existing methods can modify video content if explicit descriptions of editing targets with precise spatial locations and temporal boundaries are provided, these requirements become impractical when users attempt to conceptualize edits through implicit queries referencing semantic properties or object relationships. We introduce reasoning video editing, a task where video editing models must interpret implicit queries through multi-hop reasoning to infer editing targets before executing modifications, and a first model attempting to solve this complex task, RIVER (Reasoning-based Implicit Video Editor). RIVER decouples reasoning from generation through digital twin representations of video content that preserve spatial relationships, temporal trajectories, and semantic attributes. A large language model then processes this representation jointly with the implicit query, performing multi-hop reasoning to determine modifications, then outputs structured instructions that guide a diffusion-based editor to execute pixel-level changes. RIVER training uses reinforcement learning with rewards that evaluate reasoning accuracy and generation quality. Finally, we introduce RVEBenchmark, a benchmark of 100 videos with 519 implicit queries spanning three levels and categories of reasoning complexity specifically for reasoning video editing. RIVER demonstrates best performance on the proposed RVEBenchmark and also achieves state-of-the-art performance on two additional video editing benchmarks (VegGIE and FiVE), where it surpasses six baseline methods.

视频编辑推理模型数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。