arXiv:2512.01852cs.CLcs.AI2025-12中稿 · BHASHA Workshop @ …被引 2

首个覆盖多语言的印度幻觉识别基准,助力低资源语言模型可信度评估。

BHRAM-IL: A Benchmark for Hallucination Recognition and Assessment in Multiple Indian Languages

  • 构建涵盖9类任务的3.6万条多语言问答数据集,覆盖印地语等5种语言。
  • 14个主流多语言模型在1万条数据上测试,综合幻觉率0.23,语言修正后得分为0.385。
  • 提供可复用数据与代码,推动跨语言幻觉检测研究,适合多语言AI安全方向研究者。

大语言模型在多语言应用中日益普及,但常生成看似合理实则错误或误导性内容,即幻觉。尽管英语幻觉检测已有深入研究,但资源匮乏的印度语言仍基本未被探索。本文提出BHRAM-IL,一个覆盖印地语、古吉拉特语、马拉地语、奥里亚语及英语的多语言幻觉识别与评估基准。该基准包含36,047条精心筛选的问题,涵盖事实性、数值、推理和语言学等九类任务。我们在10,265条子集上评估了14个最先进的多语言大模型,使用类别特定指标(归一化至(0,1)区间)分析跨语言与事实性幻觉,覆盖语言、模型、规模、类别与领域。所有类别与模型汇总后,主评分为0.23,语言校正模糊得分为0.385,证明BHRAM-IL在幻觉评估中的有效性。数据集及生成与评估代码已公开于GitHub(https://github.com/sambhashana/BHRAM-IL/)与HuggingFace(https://huggingface.co/datasets/sambhashana/BHRAM-IL/),以支持未来多语言幻觉检测与缓解研究。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in multilingual applications but often generate plausible yet incorrect or misleading outputs, known as hallucinations. While hallucination detection has been studied extensively in English, under-resourced Indian languages remain largely unexplored. We present BHRAM-IL, a benchmark for hallucination recognition and assessment in multiple Indian languages, covering Hindi, Gujarati, Marathi, Odia, along with English. The benchmark comprises 36,047 curated questions across nine categories spanning factual, numerical, reasoning, and linguistic tasks. We evaluate 14 state-of-the-art multilingual LLMs on a benchmark subset of 10,265 questions, analyzing cross-lingual and factual hallucinations across languages, models, scales, categories, and domains using category-specific metrics normalized to (0,1) range. Aggregation over all categories and models yields a primary score of 0.23 and a language-corrected fuzzy score of 0.385, demonstrating the usefulness of BHRAM-IL for hallucination-focused evaluation. The dataset, and the code for generation and evaluation are available on GitHub (https://github.com/sambhashana/BHRAM-IL/) and HuggingFace (https://huggingface.co/datasets/sambhashana/BHRAM-IL/) to support future research in multilingual hallucination detection and mitigation.

幻觉检测多语言印度语评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。