arXiv:2606.02584cs.CLcs.AI2026-06

构建多语言习语理解基准,提升模型对隐喻语义的识别与解释能力

IdiomX A Multilingual Benchmark for Idiom Understanding, Retrieval, and Interpretation

  • 通过多阶段流水线构建大规模跨语言习语数据集
  • 覆盖1.2万+习语、19万+上下文例句,支持多语言语义对齐
  • 提出四任务统一评测框架,推动习语从识别到解释的进阶研究

习语因意义非组合性、依赖上下文且跨语言难以对齐,仍是自然语言处理的难点。现有习语资源规模小、上下文多样性不足或多语言覆盖有限。我们提出IdiomX,一个大规模多语言习语理解、检索与解释基准,通过可复现的多阶段流程(词典提取、大规模归一化、大模型可控增强、结构化验证)构建。数据集包含超过19万条上下文例句,涵盖1.2万+习语,提供英语、阿拉伯语、法语的语义表示对齐,以及习语/字面用法标签和丰富语言学元数据。基于此,我们定义统一的四任务基准:习语检测、上下文到习语检索、阿拉伯语到英语习语检索、习语解释,将评估从表层识别扩展至语义锚定与可解释性检索。实验表明,上下文感知的Transformer模型显著提升习语检测效果;混合检索与重排序架构大幅增强单语与跨语言检索性能。结果进一步证明,习语解释可建模为语义检索任务,引入可解释性作为补充评测维度。总体而言,IdiomX为习语研究提供了可扩展的基准,支持从检测到检索再到语义解释的演进,并具备向更多语言和隐喻推理任务拓展的模块化框架。

原文摘要 · Abstract (English)

Idiomatic expressions remain a persistent challenge for natural language processing because their meanings are often non-compositional, context-dependent, and difficult to align across languages. Existing idiom resources are often limited in scale, contextual diversity, or multilingual coverage, restricting their utility for modern language models. We introduce IdiomX, a large-scale multilingual benchmark for idiom understanding, retrieval, and interpretation, constructed through a reproducible multi-stage pipeline combining lexical resource extraction, large-scale normalization, controlled large language model enrichment, and structured validation. The resulting dataset contains over 190K contextualized examples spanning 12K+ idioms, with aligned English, Arabic, and French semantic representations, idiomatic and literal usage labels, and rich linguistic metadata. Building on this resource, we define a unified four-task benchmark covering idiom detection, context-to-idiom retrieval, Arabic-to-English idiom retrieval, and idiom interpretation, extending evaluation from figurative recognition to semantic grounding and explainable meaning retrieval. Experiments show that contextual transformer models substantially improve idiom detection, while hybrid retrieval and reranking architectures significantly strengthen both monolingual and cross-lingual idiom retrieval. Results further demonstrate that idiom interpretation can be effectively modeled as a semantic retrieval task, introducing interpretability as a complementary benchmark dimension. Overall, IdiomX provides a scalable benchmark for studying idiomatic language as a progression from detection to retrieval and semantic interpretation, and offers a modular framework extensible to additional languages and figurative reasoning tasks

习语理解多语言语义检索可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。