首个跨语言图文编辑基准,揭示多语言文本编辑中的精度退化问题
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing

- 构建3600个跨语言图文编辑样本,控制视觉变量以隔离语言影响
- 发现所有模型在希伯来语和阿拉伯语上退化最严重,文本准确性下降明显
- 提出新评估指标LSF,可捕捉字形、方向等细粒度语言错误
图文编辑已成为视觉内容创作的关键能力,但现有基准仍以英语为主,常混淆视觉合理性与语义正确性。我们提出MULTITEXTEDIT,一个包含3600个实例的受控基准,覆盖12种类型多样的语言、5个视觉领域和7种编辑操作。每个实例的语言变体共享相同视觉基础,并配有手工参考和区域掩码,使语言变量可被独立比较。为捕捉粗略文本匹配指标遗漏的书写层级错误(如缺少变音符号、左右文本顺序颠倒、混写脚本),我们引入语言保真度(LSF)指标,采用两阶段大语言模型协议评估,对母语标注者达到0.76的二次加权 extit{kappa}。使用LSF结合标准语义与掩码感知像素指标评估12个开源及专有系统,发现所有模型均存在显著跨语言退化,希伯来语和阿拉伯语最严重,荷兰语和西班牙语最小,且集中于文本准确性和书写保真度,而非整体结构。还发现普遍存在语义与像素不一致现象:输出保留全局布局和背景一致性,却扭曲特定脚本形式。
原文摘要 · Abstract (English)
Text-in-image editing has become a key capability for visual content creation, yet existing benchmarks remain overwhelmingly English-centric and often conflate visual plausibility with semantic correctness. We introduce MULTITEXTEDIT, a controlled benchmark of 3,600 instances spanning 12 typologically diverse languages, 5 visual domains, and 7 editing operations. Language variants of each instance share a common visual base and are paired with a human-edited reference and region masks, isolating the language variable for cross-lingual comparison. To capture script-level errors that coarse text-matching metrics miss, such as missing diacritics, reversed RTL order, and mixed-script renderings, we introduce a language fidelity (LSF) metric scored by a two-stage LVM protocol that first traces the edited target text and then judges it in isolation, reaching a quadratic-weighted \k{appa} of 0.76 against native-speaker annotators. Evaluating 12 open-source and proprietary systems with LSF alongside standard semantic and mask-aware pixel metrics, we find pronounced cross-lingual degradation for every model, largest on Hebrew and Arabic and smallest on Dutch and Spanish, and concentrated in text accuracy and script fidelity rather than in coarse structural dimensions. We also uncover a pervasive semantic and pixel mismatch, where outputs preserve global layout and background fidelity yet distort script-specific forms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。