arXiv:2511.09977cs.CV2025-11中稿 · AAAI被引 1

让低资源语言文本在真实图像中可编辑,还解决了风格保持难的问题。

STELLAR: Scene Text Editor for Low-Resource Languages and Real-World Data

  • 用自适应字形编码器和分阶段训练,支持多语言图文编辑。
  • 在新数据集上比现有模型视觉一致性提升2.2%,识别更准确。
  • 提出新评估指标TAS,独立衡量字体、颜色、背景相似性。

场景文本编辑(STE)旨在修改图像中的文字内容,同时保留其字体、颜色和背景等视觉风格。尽管近期基于扩散模型的方法在视觉质量上有所提升,但仍存在三大挑战:对低资源语言支持不足、合成数据与真实数据间的域差距,以及缺乏有效的风格保持评估指标。为此,我们提出STELLAR(面向低资源语言与真实数据的场景文本编辑器)。该方法通过语言自适应字形编码器和多阶段训练策略实现可靠多语言编辑,先在合成数据上预训练,再在真实图像上微调。我们还构建了新数据集STIPLAR(低资源语言与真实场景文本图像对),用于训练与评估。此外,我们提出文本外观相似性(TAS)这一新指标,通过独立测量字体、颜色和背景相似性,实现无需真实标签的鲁棒评估。实验表明,STELLAR在视觉一致性和识别准确率上均优于当前最优模型,在跨语言平均TAS上相较基线提升2.2%。

原文摘要 · Abstract (English)

Scene Text Editing (STE) is the task of modifying text content in an image while preserving its visual style, such as font, color, and background. While recent diffusion-based approaches have shown improvements in visual quality, key limitations remain: lack of support for low-resource languages, domain gap between synthetic and real data, and the absence of appropriate metrics for evaluating text style preservation. To address these challenges, we propose STELLAR (Scene Text Editor for Low-resource LAnguages and Real-world data). STELLAR enables reliable multilingual editing through a language-adaptive glyph encoder and a multi-stage training strategy that first pre-trains on synthetic data and then fine-tunes on real images. We also construct a new dataset, STIPLAR(Scene Text Image Pairs of Low-resource lAnguages and Real-world data), for training and evaluation. Furthermore, we propose Text Appearance Similarity (TAS), a novel metric that assesses style preservation by independently measuring font, color, and background similarity, enabling robust evaluation even without ground truth. Experimental results demonstrate that STELLAR outperforms state-of-the-art models in visual consistency and recognition accuracy, achieving an average TAS improvement of 2.2% across languages over the baselines.

文本编辑低资源语言图像生成评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。