测试大模型能否理解俚语并找到其真实含义
IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

- 构建跨俚语与字面表达的语义检索基准
- 现有模型在表面形式差异下匹配含义准确率低
- 适合研究语义理解与跨表达推理的学者
俚语的意义无法仅从字面推断,需超越词汇重叠的语义抽象。我们提出IdioLink,一个检索基准,用于检验模型是否能将俚语表达与概念等价的字面或改写表达关联起来。该基准包含10,700篇文档和2,140个查询,涵盖107个既有字面又有隐喻用法的俚语,每篇文档和查询均标注核心意义片段。评估BGE、E5、Contriever和Qwen等强嵌入基线模型发现,当前模型在表面形式差异较大的情况下难以准确检索等价语义,主要依赖主题和浅层语义线索。IdioLink揭示了现有模型在俚语感知语义检索中的关键缺陷,并为未来模型提供了挑战性测试平台。
原文摘要 · Abstract (English)
Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, therefore, requires semantic abstraction beyond lexical overlap. We introduce IdioLink, a retrieval benchmark designed to test whether models can link idiomatic expressions to conceptually equivalent meanings expressed in literal or paraphrased forms. IdioLink comprises 10,700 documents and 2,140 queries, spanning 107 idioms with both literal and figurative uses. Each document and query is annotated with spans that convey the core meaning. Evaluating strong embedding baselines (e.g., BGE, E5, Contriever, and Qwen), we show that current models struggle to retrieve equivalent meanings across divergent surface realizations, relying instead on topical and shallow semantic cues. IdioLink exposes key gaps in idiom-aware semantic retrieval and provides a challenging testbed for future models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。