arXiv:2512.20324cs.CL2025-12

评测大模型解孟加拉语谜题能力,发现差距显著

Can LLMs Solve My Grandma's Riddle? Evaluating Multilingual Large Language Models on Reasoning Traditional Bangla Tricky Riddles

  • 构建1244个孟加拉传统谜题的多任务评估集
  • 模型问答正确率仅56%,远低于人类83%基准
  • 适合研究低资源文化推理与模型解释性

大型语言模型在多项NLP基准上表现优异,但在隐喻性、文化背景强且资源稀缺场景下的推理能力仍待深入探索。针对孟加拉语,本文提出BanglaRiddleEval,一个包含1,244个传统孟加拉语谜题的基准,涵盖四个任务(共4,976个谜题-任务实例)。通过基于LLM的流程,生成链式思维解释、语义连贯干扰项及细粒度歧义标注,并在不同提示策略下评估多种开源与闭源模型。结果显示,生成式问答的语义重叠中等,但正确率低;多项选择题准确率最高约56%,而人类基准为83%;歧义解析准确率在26%至68%之间,高质量解释仅限最强模型。这表明当前模型虽能捕捉部分线索,但仍远未达到人类水平,验证了BanglaRiddleEval作为低资源隐喻推理新基准的挑战性。所有数据、代码与评估脚本已公开于GitHub:https://github.com/Labib1610/BanglaRiddleEval。

原文摘要 · Abstract (English)

Large Language Models (LLMs) show impressive performance on many NLP benchmarks, yet their ability to reason in figurative, culturally grounded, and low-resource settings remains underexplored. We address this gap for Bangla by introducing BanglaRiddleEval, a benchmark of 1,244 traditional Bangla riddles instantiated across four tasks (4,976 riddle-task artifacts in total). Using an LLM-based pipeline, we generate Chain-of-Thought explanations, semantically coherent distractors, and fine-grained ambiguity annotations, and evaluate a diverse suite of open-source and closed-source models under different prompting strategies. Models achieve moderate semantic overlap on generative QA but low correctness, MCQ accuracy peaks at only about 56% versus an 83% human baseline, and ambiguity resolution ranges from roughly 26% to 68%, with high-quality explanations confined to the strongest models. These results show that current LLMs capture some cues needed for Bangla riddle reasoning but remain far from human-level performance, establishing BanglaRiddleEval as a challenging new benchmark for low-resource figurative reasoning. All data, code, and evaluation scripts are available on GitHub: https://github.com/Labib1610/BanglaRiddleEval.

大模型评测低资源语言文化推理谜题理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。