arXiv:2602.12921cs.CL2026-02中稿 · presentation at LR…被引 2

构建首个大规模孟加拉语习语数据集,揭示大模型在隐喻理解上的显著不足。

When Words Don't Mean What They Say: Figurative Understanding in Bengali Idioms

  • 构建10,361条孟加拉语习语的19维标注数据集,涵盖语义、文化等维度
  • 30个主流模型在习语理解任务中准确率均未超50%,远低于人类83.4%
  • 为低资源语言的隐喻理解与文化语境建模提供基础工具和基准

隐喻语言理解仍是大型语言模型(LLMs)的重大挑战,尤其对低资源语言而言。为此,我们引入一个新习语数据集——包含10,361条孟加拉语习语的大规模、文化根基型语料库。每条习语均经专家共识过程建立并优化的19维标注体系进行标注,涵盖其语义、句法、文化及宗教维度,为计算语言学提供丰富结构化资源。为建立可靠的孟加拉语隐喻理解基准,我们在推断习语含义的任务上评估了30个最先进的多语言与指令微调模型。结果揭示出显著性能差距:无一模型准确率超过50%,远低于人类83.4%的表现,凸显现有模型在跨语言与文化推理方面的局限性。通过发布该习语数据集与基准,我们为提升孟加拉语及其他低资源语言的隐喻理解与文化语境建模奠定了基础框架。

原文摘要 · Abstract (English)

Figurative language understanding remains a significant challenge for Large Language Models (LLMs), especially for low-resource languages. To address this, we introduce a new idiom dataset, a large-scale, culturally-grounded corpus of 10,361 Bengali idioms. Each idiom is annotated under a comprehensive 19-field schema, established and refined through a deliberative expert consensus process, that captures its semantic, syntactic, cultural, and religious dimensions, providing a rich, structured resource for computational linguistics. To establish a robust benchmark for Bangla figurative language understanding, we evaluate 30 state-of-the-art multilingual and instruction-tuned LLMs on the task of inferring figurative meaning. Our results reveal a critical performance gap, with no model surpassing 50% accuracy, a stark contrast to significantly higher human performance (83.4%). This underscores the limitations of existing models in cross-linguistic and cultural reasoning. By releasing the new idiom dataset and benchmark, we provide foundational infrastructure for advancing figurative language understanding and cultural grounding in LLMs for Bengali and other low-resource languages.

隐喻理解低资源语言文化语境习语数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。