arXiv:2606.04479cs.CVcs.AI2026-06中稿 · CVPR

测试图像中文本推理的忠实度,发现现有效果看似清晰实则常出错。

Evaluating Reasoning Fidelity in Visual Text Generation

论文配图:Evaluating Reasoning Fidelity in Visual Text Generation
图 1 · 摘自论文原文
  • 用图像生成方式表达完整推理过程来评估模型可靠性。
  • 模型在长文本、事实判断、上下文理解中频繁出现语义错误和逻辑不一致。
  • 适合关注视觉推理可信度的研究者和应用开发者。

近期文本到图像(T2I)模型能在图像中生成高度可读且结构良好的文本,适用于文档生成与幻灯片制作等场景。然而,当复杂解题过程需直接通过渲染文本表达时,这些系统是否真正保留了推理能力仍不明确,抑或仅模仿表面模式?我们通过评估视觉文本生成中的推理忠实度来探究此问题,要求模型将完整的推理过程以图像形式呈现。评估涵盖长文本渲染、事实知识探测、上下文理解及多步推理。在各项任务中,当前T2I模型尽管生成的文本视觉清晰,却频繁出现语义错误、逻辑矛盾和中间步骤错误。这些缺陷与文本模型在同一任务上的优异表现形成鲜明对比。研究揭示了视觉文本生成与程序化推理之间存在显著差距,呼吁发展更可靠的视觉文本推理能力。

原文摘要 · Abstract (English)

Recent text-to-image (T2I) models can render highly legible and well-structured text within images, enabling applications including document generation and slide generation. However, it remains unclear whether such systems faithfully preserve reasoning ability when complex solutions must be expressed directly through rendered text, or whether they merely imitate surface-level patterns. We investigate this question by evaluating reasoning fidelity in visual text generation, where models must express complete reasoning processes as images. Our evaluation includes long text rendering, factual knowledge probing, context understanding, and multi-step reasoning. Across these settings, we find that current T2I models frequently produce semantic errors, logical inconsistencies, and incorrect intermediate steps, even when the rendered text appears visually clear. These failures contrast with the strong reasoning performance of text-only models on the same tasks. Our findings reveal a substantial gap between visual text generation and procedural reasoning, motivating more reliable visual text reasoning.

文本生成视觉推理模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。