arXiv:2511.20399cs.CLcs.AI2025-11被引 1

构建首个孟加拉语隐喻与文化推理低资源数据集,揭示大模型在文化理解上的短板。

BengaliFig: A Low-Resource Challenge for Figurative and Culturally Grounded Reasoning in Bengali

  • 收集435个孟加拉民间文学谜题,多维标注推理类型、文化深度等五项属性。
  • 八款主流大模型在零样本和少样本下均暴露出隐喻与文化特异性推理缺陷。
  • 适合关注低资源语言、文化敏感性评估与包容性NLP的研究者使用。

大型语言模型在广泛多语言基准上表现优异,但在隐喻与文化背景推理方面,尤其在低资源语境下,仍缺乏充分评估。我们提出BengaliFig,一个紧凑但标注丰富的挑战数据集,聚焦孟加拉语这一广泛使用却资源匮乏的语言。该数据集包含435个源自孟加拉口头与文学传统的独特谜题,每条数据在五个正交维度上进行标注:推理类型、陷阱类型、文化深度、答案类别与难度,并通过约束感知的AI辅助流水线自动转为多项选择格式。我们在零样本与少样本链式思维提示下,评估了来自主要厂商的八款前沿大模型,发现其在隐喻与文化特定推理方面存在持续性弱点。BengaliFig不仅为评估大模型在低资源文化语境下的鲁棒性提供了诊断工具,也推动了更具包容性与遗产意识的自然语言处理评估发展。

原文摘要 · Abstract (English)

Large language models excel on broad multilingual benchmarks but remain to be evaluated extensively in figurative and culturally grounded reasoning, especially in low-resource contexts. We present BengaliFig, a compact yet richly annotated challenge set that targets this gap in Bengali, a widely spoken low-resourced language. The dataset contains 435 unique riddles drawn from Bengali oral and literary traditions. Each item is annotated along five orthogonal dimensions capturing reasoning type, trap type, cultural depth, answer category, and difficulty, and is automatically converted to multiple-choice format through a constraint-aware, AI-assisted pipeline. We evaluate eight frontier LLMs from major providers under zero-shot and few-shot chain-of-thought prompting, revealing consistent weaknesses in metaphorical and culturally specific reasoning. BengaliFig thus contributes both a diagnostic probe for evaluating LLM robustness in low-resource cultural contexts and a step toward inclusive and heritage-aware NLP evaluation.

低资源文化推理大模型评估孟加拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。