评测图像中文本编辑的推理能力,揭示模型在语义与布局一致性上的短板。
TextEditBench: Evaluating Reasoning-aware Text Editing Beyond Rendering
- 构建聚焦文本区域的综合评测基准,强调语义与跨模态推理。
- 现有模型在上下文依赖和物理合理性上表现不佳,准确率不足60%。
- 适合研究多模态生成、文本编辑与推理的学者使用。
文本渲染已成为视觉生成领域最具挑战性的前沿之一,受到大规模扩散模型和多模态模型的广泛关注。然而,图像中的文本编辑仍处于探索阶段,因其需在保持语义、几何和上下文连贯性的同时生成可读字符。为此,我们提出TextEditBench,一个专注于图像中以文本为中心区域的综合性评估基准。该基准不仅涵盖基础像素操作,更强调需要理解物理合理性、语言意义及跨模态依赖的推理密集型编辑场景。我们进一步提出新的评估维度——语义期望(Semantic Expectation, SE),用于衡量模型在文本编辑过程中维持语义一致性、上下文连贯性和跨模态对齐的能力。对当前主流编辑系统的广泛实验表明,尽管模型能遵循简单指令,但在上下文依赖推理、物理一致性与布局感知融合方面仍存在明显不足。通过聚焦这一长期被忽视但根本性的能力,TextEditBench为推进文本引导的图像编辑与多模态生成中的推理能力建立了新测试平台。
原文摘要 · Abstract (English)
Text rendering has recently emerged as one of the most challenging frontiers in visual generation, drawing significant attention from large-scale diffusion and multimodal models. However, text editing within images remains largely unexplored, as it requires generating legible characters while preserving semantic, geometric, and contextual coherence. To fill this gap, we introduce TextEditBench, a comprehensive evaluation benchmark that explicitly focuses on text-centric regions in images. Beyond basic pixel manipulations, our benchmark emphasizes reasoning-intensive editing scenarios that require models to understand physical plausibility, linguistic meaning, and cross-modal dependencies. We further propose a novel evaluation dimension, Semantic Expectation (SE), which measures reasoning ability of model to maintain semantic consistency, contextual coherence, and cross-modal alignment during text editing. Extensive experiments on state-of-the-art editing systems reveal that while current models can follow simple textual instructions, they still struggle with context-dependent reasoning, physical consistency, and layout-aware integration. By focusing evaluation on this long-overlooked yet fundamental capability, TextEditBench establishes a new testing ground for advancing text-guided image editing and reasoning in multimodal generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。