arXiv:2505.16216cs.CL2025-05EMNLP被引 13

探究大模型理解成语是靠记忆还是推理,发现其结合了记忆与上下文推理。

Memorization or Reasoning? Exploring the Idiom Understanding of LLMs

  • 构建六语言成语数据集MIDAS,评估大模型成语理解能力。
  • 模型在复合成语上表现更好,说明具备一定推理能力。
  • 适合研究多语言语义理解与模型内部机制的学者参考。

成语因其独特的语言特性长期以来构成挑战,与普通表达有明显差异。尽管近期研究已利用大语言模型(LLMs)处理多种任务中的成语,如含成语的句子生成和成语性机器翻译,但对大模型在多语言环境下成语处理机制的理解仍不充分。为此,我们引入MIDAS,一个涵盖六种语言的大型成语数据集,每条成语均配有对应释义。基于此资源,我们对大模型的成语处理能力进行了全面评估,识别出影响其性能的关键因素。研究发现,大模型不仅依赖记忆,还采用混合方法,结合上下文线索与推理,尤其在处理复合成语时表现显著。这表明,大模型的成语理解源于内部知识检索与基于推理的推断之间的交互作用。

原文摘要 · Abstract (English)

Idioms have long posed a challenge due to their unique linguistic properties, which set them apart from other common expressions. While recent studies have leveraged large language models (LLMs) to handle idioms across various tasks, e.g., idiom-containing sentence generation and idiomatic machine translation, little is known about the underlying mechanisms of idiom processing in LLMs, particularly in multilingual settings. To this end, we introduce MIDAS, a new large-scale dataset of idioms in six languages, each paired with its corresponding meaning. Leveraging this resource, we conduct a comprehensive evaluation of LLMs' idiom processing ability, identifying key factors that influence their performance. Our findings suggest that LLMs rely not only on memorization, but also adopt a hybrid approach that integrates contextual cues and reasoning, especially when processing compositional idioms. This implies that idiom understanding in LLMs emerges from an interplay between internal knowledge retrieval and reasoning-based inference.

大模型成语理解多语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。