arXiv:2609.02502cs.CVcs.AI2026-09

首个评测文生图模型生成视觉隐喻能力的基准

Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models

论文配图:Blending Concepts: Benchmarking Visual Metaphor Generation in Text-to-Image Models
图 1 · 摘自论文原文
  • 构建1500个真实创意图像组成的隐喻数据集
  • 11个主流模型在跨域组合上表现普遍不足
  • 适合关注创意生成与跨模态理解的研究者

文生图(T2I)模型在准确呈现指定物体和属性方面已取得显著进展,但其生成视觉隐喻——通过融合两个不同领域元素来表达抽象概念的图像——的能力仍缺乏系统评估。为此,我们提出VMetaphor-Bench,首个用于评估T2I模型视觉隐喻生成能力的基准。该基准包含从真实创意图像中筛选出的1500个视觉隐喻,按三个层级和十类主题组织,每条样本配有两条不同具体程度的提示。评估采用混合框架,在多模态大模型作为裁判的范式下,结合9594个问题的多选题协议(涵盖四个隐喻保真度层级)和基于三个感知维度的评分协议。对11个代表性T2I模型的广泛测试表明,即使最强的专有模型在组合结构构建和跨域映射方面仍存在明显短板,凸显视觉隐喻生成是未来T2I研究的重要前沿。

原文摘要 · Abstract (English)

Text-to-image (T2I) models have achieved remarkable success at faithfully rendering specified objects and attributes, yet their ability to produce visual metaphors, images that convey abstract ideas by combining elements from two distinct domains, remains largely unexamined. To bridge this gap, we introduce VMetaphor-Bench, the first benchmark for evaluating visual metaphor generation in T2I models. It comprises 1,500 visual metaphors curated from real-world creative imagery, organized into three levels and ten categories, with each sample paired with two prompts of differing specificity. For evaluation, we develop a hybrid framework within an MLLM-as-judge paradigm, combining a multiple-choice question (MCQ) based protocol of 9,594 questions across four levels of metaphorical fidelity with a dimension-based scoring protocol along three perceptual dimensions. Extensive evaluation of 11 representative T2I models reveals that even the strongest proprietary models struggle with compositional structuring and cross-domain mapping, key aspects of metaphorical expression, highlighting visual metaphor generation as an important frontier for future T2I research.

视觉隐喻文生图评估基准跨域生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。