构建首个文本主导图像编辑数据集与评测框架,提升多语言文字修改精度。
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing
- 基于HTML自动生成33万组图文编辑训练对,覆盖15种语言
- 提出两阶段训练:字形引导微调+多目标强化学习,显著提升文字清晰度
- 专为复杂文字操作设计,适合多语言内容创作与AI绘图工具开发者
基于指令的图像编辑旨在根据用户指令修改图像中的特定内容,同时保留非目标区域。除传统对象与风格编辑外,文本主导编辑聚焦于修改、翻译或重排图像中嵌入的文字元素。然而现有模型在执行复杂文本编辑时常出现模糊或幻觉字符。我们归因于缺乏针对该任务的专用训练范式,以及缺少大规模数据集与标准化评测基准。为此,我们提出WeEdit,涵盖可扩展的数据构建流程、两个评测基准及定制化的两阶段训练策略。具体而言,我们设计了一种新型基于HTML的自动编辑流程,生成33万组训练样本,覆盖多样化编辑操作和15种语言,并提供标准双语与多语言评测基准。算法上,采用字形引导的监督微调注入显式空间与内容先验,再通过多目标强化学习阶段对齐指令遵循性、文字清晰度与背景保留。大量实验表明,WeEdit在多种编辑操作下显著优于现有开源模型。
原文摘要 · Abstract (English)
Instruction-based image editing aims to modify specific content within existing images according to user-provided instructions while preserving non-target regions. Beyond traditional object- and style-centric manipulation, text-centric image editing focuses on modifying, translating, or rearranging textual elements embedded within images. However, existing leading models often struggle to execute complex text editing precisely, frequently producing blurry or hallucinated characters. We attribute these failures primarily to the lack of specialized training paradigms tailored for text-centric editing, as well as the absence of large-scale datasets and standardized benchmarks necessary for a closed-loop training and evaluation system. To address these limitations, we present WeEdit, a systematic solution encompassing a scalable data construction pipeline, two benchmarks, and a tailored two-stage training strategy. Specifically, we propose a novel HTML-based automatic editing pipeline, which generates 330K training pairs covering diverse editing operations and 15 languages, accompanied by standardized bilingual and multilingual benchmarks for comprehensive evaluation. On the algorithmic side, we employ glyph-guided supervised fine-tuning to inject explicit spatial and content priors, followed by a multi-objective reinforcement learning stage to align generation with instruction adherence, text clarity, and background preservation. Extensive experiments demonstrate that WeEdit outperforms previous open-source models by a clear margin across diverse editing operations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。