arXiv:2609.03410cs.CL2026-09

首个孟加拉语习语数据集,评测大模型理解能力

To What Extent Do Large Language Models Understand Bangla Idioms?

  • 构建首个大规模孟加拉语习语基准数据集
  • 不同模型在习语识别任务中表现差异显著
  • 适合低资源语言习语研究者参考

习语是自然语言的重要组成部分,反映文化特征,对计算模型构成独特挑战,尤其在低资源语言中。本文首次提出大规模孟加拉语习语基准数据集,并配套生成用于习语含义识别的合成多选题数据集。我们通过零样本与少样本提示策略,全面评估近期大语言模型(LLMs)在三个习语相关任务上的表现:释义、习语片段检测和含义识别。结果表明模型性能存在显著差异,无单一模型在所有任务上持续领先。值得注意的是,Phi-4-mini-instruct 在释义任务中表现最佳,Kimi-K2-32b-instruct 在片段检测中领先,Gemini-2.5-flash 在含义识别中表现最优。我们认为该数据集与分析将为提升大模型对习语的理解能力,尤其是孟加拉语及其他低资源语言的研究提供重要支持。

原文摘要 · Abstract (English)

Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms, complemented by a synthetic multiple-choice question (MCQ) dataset for idiom meaning identification. We conduct a comprehensive evaluation of recent large language models (LLMs) across three idiom-related tasks: paraphrasing, idiom span detection, and meaning identification, leveraging zero-shot and few-shot prompting strategies. Our results reveal substantial variability in model performance, with no single LLM consistently outperforming others across all tasks. Notably, Phi-4-mini-instruct excels in paraphrasing, Kimi-K2-32b-instruct in span detection, and Gemini-2.5-flash in meaning identification. We believe that our datasets and analyses will provide valuable resources to guide future research in improving LLM comprehension of idiomatic expressions, particularly in Bangla and other low-resource languages.

习语理解低资源语言大模型评测孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。