用思维链引导视频编辑,让文字指令更精准地控制物体位置和动作。
CoT-Edit: Let CoT Guide Instruction Video Editing

- 通过思维链增强的多模态大模型生成带边框和属性的编辑指令
- 在复杂场景中实现多相似物体的精准定位与物理合理新增
- 适合需要高精度视频编辑的应用,如影视制作、虚拟演播
基于文本的指令式视频编辑在复杂场景中仍具挑战:纯文本提示常无法捕捉精确的空间关系与物理约束,导致目标模糊和物理上不合理的结果。为此,我们提出一种计划-引导-编辑框架,显式连接语义意图与空间执行。框架中,一个增强思维链的多模态大语言模型(MLLM)作为规划器,对视频与指令进行结构化推理,生成精确的边界框序列与属性丰富的编辑指令。这些空间先验引导一个框条件掩码生成器,将模糊的全局检索转化为局部化、上下文感知的优化,生成更准确捕捉物体尺度、接触关系与位置的掩码。结合这些空间与语义信号,基于扩散模型的编辑器融合掩码、增强指令与帧特征,生成高保真度且时空一致的编辑结果。采用分模块训练后联合优化,本框架在减少数据需求的同时实现更优性能,在多个相似物体场景中完成精准定位与物理一致的物体添加,大量实验表明其优于多种强基线方法。
原文摘要 · Abstract (English)
Text-driven instruction-based video editing in complex scenes remains challenging: purely textual prompts often fail to capture precise spatial relationships and physical constraints, resulting in target ambiguity and physically implausible outcomes. To address this, we propose a plan--guide--edit framework that explicitly bridges semantic intent and spatial execution. In our framework, a Chain-of-Thought (CoT)-enhanced multimodal large language model (MLLM) serves as a planner, performing structured reasoning over the video and instructions to derive a precise sequence of bounding boxes and attribute-enriched editing directives. These spatial priors then guide a box-conditioned mask generator, transforming ambiguous global retrieval into localized, context-aware refinement and producing masks that more accurately capture object scale, contact relationships, and placement. Building on these spatial and semantic signals, a diffusion-based editor integrates the masks, enriched instructions, and frame features to render high-fidelity edits that remain temporally coherent and spatially well aligned. Trained first in a modular manner and then jointly, our framework achieves superior performance with reduced data requirements, delivering precise localization in scenes with multiple similar objects and physically consistent object additions, and extensive experiments demonstrate state-of-the-art performance over multiple strong baseline methods. More details are available at: https://github.com/flying-sky999/CoT-Edit
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。