arXiv:2608.26809cs.CVcs.MM2026-08

用智能体推理实现长视频多指令编辑,保持一致性与时空连贯性。

Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

论文配图:Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning
图 1 · 摘自论文原文
  • 构建智能体框架,融合大语言模型与视觉语言模型解析多指令
  • 在真实动态场景中实现无幻觉、跨片段一致的编辑结果
  • 适合需要精准长视频批量处理的研究者与创作者

尽管生成式AI已显著推进视频编辑技术,现有方法主要聚焦单次或短片段编辑。对包含多个指令的长视频进行编辑仍面临巨大挑战。简单的分段策略(如固定时长分割)常导致实体碎片化、严重编辑幻觉和时间连续性破坏。为此,我们提出多指令多片段长视频编辑(MMLVE)任务,围绕三大目标:跨片段编辑一致性(CSEC)、多指令解耦(MID)和时空结构零破坏(ZDSS)。为此,我们设计了一个智能体编辑框架,利用大语言模型(LLMs)与视觉语言模型(VLMs)的协同作用,实现片段级视频解耦与精准指令解析。为全面评估该任务,我们构建了MMLVE-Bench数据集,其特点为复杂真实时空动态、高密度异构指令及稀疏随机实体分布。进一步引入三项针对MMLVE的评估指标。大量实验表明,我们的MMLVE-Agent优于现有闭源最先进方法(如Seedance 2.0),成功消除编辑幻觉,保持跨片段一致性,并实现平滑的时空过渡。

原文摘要 · Abstract (English)

While generative AI has significantly advanced video editing, existing methods primarily focus on single-shot or short video clips. Editing long videos with multiple instructions remains a formidable challenge. Naive chunking strategies, e.g., fixed-duration segmentation, often lead to entity fragmentation, severe editing hallucinations, and disrupted temporal continuity. To bridge this gap, we introduce the Multi-Instruction Multi-Shot Long-Video Editing (MMLVE) task, which is structured around three core objectives: Cross-Shot Editing Consistency (CSEC), Multi-Instruction Decoupling (MID), and Zero-Destruction on Spatiotemporal Structure (ZDSS). To tackle these three unique challenges, we introduce an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing. Furthermore, to comprehensively evaluate this task, we construct MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions. Three MMLVE-focused evaluation metrics are further exploited to assess the quality of the editing results. Extensive experiments demonstrate that our MMLVE-Agent outperforms existing closed-source SOTA approaches (e.g., Seedance 2.0), successfully eliminating editing hallucinations, preserving cross-shot editing consistency, and attaining seamless spatiotemporal transitions.

视频编辑智能体多指令长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。