研究多模态模型如何理解多义词,发现图像生成比文本更单一。
Where did the ambiguity go? Examining how multimodal models interpret polysemous words

- 用无上下文的多义词测试17个文生图和15个文生模型
- 图像生成的语义熵仅为0.10,远低于人类想象的0.47
- 模型预测的语义分布比实际输出更丰富,暴露理解偏差
人类语言高度多义,许多常见词(如'bank'或'palm')具有多个含义,影响沟通与想象。大型语言模型(LLMs)已显示出对这种语义多样性的理解,但对图像等其他模态中多义性如何呈现仍知之甚少。我们通过向17个文生图模型和15个文生模型输入无上下文的多义词,测量其在多次生成中产生的语义分布。结果表明存在明显的多模态差距:在每个模型家族中,生成图像所覆盖的语义远少于生成文本(归一化熵0.10对比0.25),且两者均远低于人类对同一词语的想象(归一化熵0.47)。然而,当要求模型列出各可能含义的生成频率时,其预测分布反而比实际输出更多样化。这些结果揭示了基础模型在跨模态表达意义上的差距,以及其理解在不同模态间无法忠实传递的问题。
原文摘要 · Abstract (English)
Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。