arXiv:2608.04750cs.CVcs.CL2026-08中稿 · as a full paper at…

测试文生图模型对隐喻的理解能力,发现它们常把比喻对象当实物生成。

Simile Understanding in Text-to-Image Models: An Evaluation Framework

论文配图:Simile Understanding in Text-to-Image Models: An Evaluation Framework
图 1 · 摘自论文原文
  • 构建可控隐喻数据集,用可检测物体作比喻载体
  • 用YOLO检测器自动评估生成图像中比喻物是否存在
  • 分析扩散模型中间层,追踪比喻物如何被逐步生成

隐喻以简洁而富有表现力的方式描述视觉特征。近期文生图模型(t2i models)能从隐喻提示生成视觉上吸引人的图像,但即使顶尖模型也频繁误解隐喻的载体,并将其与本体混淆。这些系统性错误揭示了文生图模型在修辞语言与实体视觉定位之间的差距。为此,我们提出一种可扩展的隐喻理解评估框架:(1) 构建一个受控的隐喻数据集,其中比喻载体来自预定义的可检测物体类别,并与多样模板组合;(2) 基于YOLO(You Only Look Once)检测的自动定位度量;(3) 使用Diffusion Lens分析文本编码器层,追踪生成过程中比喻载体的出现过程。在多种架构的t2i模型上进行实验,均发现一致的字面化失败模式。我们进一步讨论了提升隐喻视觉定位能力的潜在策略。

原文摘要 · Abstract (English)

Similes provide a compact and expressive way to describe visual characteristics in text prompts. Recent text-to-image models (t2i models) can produce visually compelling outputs from simile prompts, yet even frontier models frequently misinterpret the metaphorical vehicle and confuse it with the object. These systematic failures reveal a gap between figurative language and object-level visual grounding in t2i models. To investigate this issue, we propose a scalable evaluation framework for simile understanding. Our framework includes (1) a controlled simile dataset in which metaphorical vehicles are drawn from a predefined set of object-detectable categories and combined with diverse templates, (2) automatic grounding metrics based on YOLO (You Only Look Once) detection, and (3) text encoder layer analysis using Diffusion Lens to track how metaphorical vehicles emerge during generation. Experiments across architecturally diverse t2i models reveal consistent literalization failure patterns. We further discuss potential mitigation strategies for improving simile grounding in t2i models.

隐喻理解文生图评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。