首个针对孟加拉语大模型幻觉的多任务评估框架,揭示了现有方法在低资源语言中的不足。
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
- 构建涵盖4项任务的细粒度幻觉评估框架,生成1.2万条幻觉样本。
- 7个模型在双轨评测中得分介于7.72%至55.42%,显示幻觉校准差异显著。
- 提出双轨校准指标BenHalluScore,证明仅靠提示工程无法可靠降低幻觉。
尽管孟加拉语是全球第六大使用语言,但此前尚无系统性研究评估其大语言模型(LLMs)的幻觉问题。本文提出BenHalluEval,一个覆盖生成式问答(GQA)、孟加拉-英语混用问答、摘要和推理四类任务的细粒度幻觉评估框架。利用GPT-5.4构建12,000条幻觉候选样本,涵盖十二种任务特定类型,数据源自三个现有孟加拉语数据集。对七种不同类别(推理导向、多语言、孟加拉语专精)的LLM进行双轨评估:轨道A衡量真实样本上的假阳性率,轨道B评估幻觉样本检测率。为同时惩罚两类错误并防止统一响应偏差导致分数虚高,提出BenHalluScore双轨校准指标,模型与任务间得分范围为7.72%至55.42%,揭示幻觉校准存在显著差异。链式思维提示虽改变响应分布,但未一致提升幻觉识别能力。BenHalluEval建立了首个面向孟加拉语的幻觉基准,凸显单一轨道及仅靠提示工程在低资源语言场景下的局限性。数据集与代码已公开于https://anonymous.4open.science/r/BanglaHalluEval-EB77。
原文摘要 · Abstract (English)
Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali. We introduce BenHalluEval, a fine-grained hallucination evaluation framework for Bengali covering four tasks: Generative Question Answering (GQA), Bangla-English Code-Mixed QA, Summarization, and Reasoning. We construct 12,000 hallucinated candidates using GPT-5.4 across twelve task-specific hallucination types, drawn from three existing Bengali datasets, and evaluate seven LLMs spanning reasoning-oriented, multilingual, and Bengali-centric categories under a dual-track protocol that independently measures false-positive rate on ground-truth instances (Track A) and hallucination detection rate on hallucinated candidates (Track B). To jointly penalise both failure modes and prevent inflated scores from uniform response bias, we propose BenHalluScore, a dual-track calibration metric that ranges from 7.72% to 55.42% across models and tasks, revealing substantial variation in hallucination calibration. Chain-of-thought prompting, applied as a mitigation strategy, shifts response distributions without consistently improving hallucination discrimination. BenHalluEval establishes the first dedicated hallucination benchmark for Bengali and highlights the inadequacy of single-track and prompting-only evaluation approaches for low-resource language settings. The dataset and code are available at https://anonymous.4open.science/r/BanglaHalluEval-EB77.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。