融合文字与视觉指令,实现精准又符合意图的图像编辑。
Text-Vision Co-Instructed Image Editing

- 用文字表达意图,用拖拽/点击提供位置,双模态协同控制。
- 在23000+样本数据上训练,编辑精度和语义一致性显著提升。
- 适合需要精确控制且保持原图结构的研究者与设计师。
现有图像编辑方法分为基于文本指令和基于视觉提示两类。文本指令语义丰富,但空间控制粗糙;视觉提示(如拖拽、点击)可精确定位,但语义意图模糊。为此,本文提出文本-视觉联合指导的图像编辑框架,将文本作为语义意图,稀疏视觉指令作为空间引导,实现精准且忠实于意图的编辑。研究构建了一个包含超过23,000个样本的动态视频衍生的图文-视觉指令配对数据集,支持跨模态对齐监督。提出TV-Edit框架,通过上下文建模将视觉指令转化为语义感知的控制表示,输入预训练编辑模型。结合语义与空间约束后,相比纯文本或仅拖拽方案,显著提升空间控制精度、降低指令歧义性,并增强结构一致性。此外,建立TV-Edit-Bench基准,通过真实参考图像和受控的文本-视觉变化,评估语义忠实度、空间对齐性和视觉一致性。多模型实验表明,TV-Edit始终优于当前最优的指令式与拖拽式基线。
原文摘要 · Abstract (English)
Existing image editing methods can be generally categorized into textual instruction-based and visual prompt-based ones. Textual instructions are semantically expressive, but are limited by the coarse granularity of spatial control of the editing results. In contrast, visual prompts such as drag and point can provide precise spatial guidance, but are limited by the inherent ambiguity in semantic intent. To unify the strength of textual and visual prompts, we present Text-Vision Co-Instructed Image Editing, which jointly models textual instructions as semantic intent and sparse visual instructions as spatial guidance, aiming to achieve precise and intent-faithful image manipulation. To this end, we first construct a textual-visual instruction paired dataset with more than 23K samples derived from dynamic videos, enabling aligned supervision for cross-modal instruction. We then propose TV-Edit, a Textual-Visual instruction unified Editing framework to contextualize drag or point-based visual instructions with image-text semantics and lift them into semantic-aware control representations for pretrained editing backbones. By integrating semantic intent and spatial constraints, TV-Edit leads to more precise spatial control, less instruction ambiguity, and stronger structural consistency than text-only or drag-based alternatives. Finally, we establish TV-Edit-Bench, a deliberately designed benchmark to evaluate semantic faithfulness, spatial alignment, and visual consistency with ground-truth references and controlled textual-visual variations for reliable assessment. Our experiments across multiple editing backbones demonstrate that TV-Edit consistently yields more precise and intent-faithful edits, significantly outperforming state-of-the-art instruction-based and drag-based baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。