arXiv:2608.09097cs.CV2026-08中稿 · ACM MM 2026

用草图+指令实现像素级精准局部图像编辑

SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision

论文配图:SI-Edit: Toward Sketch-Instruction Guided Local Image Editing with Pixel-Level Precision
图 1 · 摘自论文原文
  • 结合草图与语义指令,实现空间与语义协同控制
  • 在草图对齐度上优于现有方法,支持精细局部修改
  • 适合需要精确图像编辑的研究者和设计师

尽管生成模型发展迅速,基于草图的局部图像编辑仍难以实现像素级精度,尤其在细粒度形变方面。主要瓶颈在于缺乏同时提供几何约束与语义指令的高质量公开数据集。为此,我们提出**SI-Data**,一个专为指令引导的局部草图编辑设计的高质量数据集。通过多模态大语言模型(MLLMs)构建自动化流水线,合成包含原始图像、局部几何草图、语义指令及对应编辑图像的四元组。该数据集同时提供可靠的空间锚点与明确的语义意图,支持空间-语义联合学习。在此基础上,我们提出**SI-Edit**框架,融合语义指令与精确几何约束。为弥补评估标准缺失,我们建立了一套综合指标,用于衡量结构保真度(如草图-边缘对齐)与语义一致性。实验表明,相较于基线方法,SI-Edit在草图引导图像编辑中具备更可靠的结构控制能力,可实现与用户意图一致的像素级局部精修。代码与数据已公开于[项目页](https://github.com/ywxsuperstar/SIEdit)。

原文摘要 · Abstract (English)

Despite rapid advances in generative models, achieving pixel-level precision in sketch-based image editing remains a persistent challenge, particularly for fine-grained local deformations. This gap stems primarily from the critical shortage of high-quality, publicly available benchmark datasets that jointly provide geometric constraints and semantic instructions. To address this issue, we first introduce **SI-Data**, a high-quality dataset specifically designed for instruction-guided local sketch editing. We develop an automated pipeline leveraging Multimodal Large Language Models (MLLMs) to synthesize comprehensive quadruplets comprising original images, local geometric sketches, semantic instructions, and corresponding edited images. By providing both reliable spatial anchors and explicit semantic intent, SI-Data uniquely enables collaborative spatial-semantic learning. Building upon this, we propose a collaborative framework called **SI-Edit** that integrates semantic instructions with precise geometric constraints. Furthermore, to address the lack of standardized evaluation, we establish a comprehensive set of metrics designed to measure both structural fidelity (e.g., sketch-to-edge alignment) and semantic adherence. Experimental results demonstrate that SI-Edit provides more reliable structural control than baselines for sketch-based image editing, and achieves precise, pixel-level local refinements aligned with user intent. The data and code are released on the [project page](https://github.com/ywxsuperstar/SIEdit).

图像编辑草图生成多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。