arXiv:2601.16766cs.CLcs.AI2026-01Conference of the …被引 2

低资源语言下,幻觉检测器比任务准确率更稳定。

Do LLM hallucination detectors suffer from low-resource effect?

  • 在低资源语言中测试幻觉检测器的鲁棒性
  • 检测器准确率下降幅度仅为任务准确率的几分之一
  • 跨语言检测需本地监督,适合多语言系统研究者

大型语言模型虽在诸多任务上超越人类,但仍会以不可预见的方式失败。本文聚焦两大常见故障模式:(i)幻觉,即模型生成关于世界的错误信息;(ii)低资源效应,即模型在英语等高资源语言表现优异,但在孟加拉语等低资源语言中性能显著下降。我们研究两者交集,提出问题:幻觉检测器是否也受低资源效应影响?在三个领域(事实回忆、科学与数学、人文学科)的五个任务上,对四种模型和三种检测器进行实验发现:如预期,低资源语言的任务准确率大幅下降。但检测器的准确率下降幅度通常仅为任务准确率下降的几分之一。结果表明,即使在低资源语言中,模型内部可能仍保留其不确定性的信号。检测器在同语言内及多语言设置中表现稳健,但在无本地监督的跨语言场景中失效。

原文摘要 · Abstract (English)

LLMs, while outperforming humans in a wide range of tasks, can still fail in unanticipated ways. We focus on two pervasive failure modes: (i) hallucinations, where models produce incorrect information about the world, and (ii) the low-resource effect, where the models show impressive performance in high-resource languages like English but the performance degrades significantly in low-resource languages like Bengali. We study the intersection of these issues and ask: do hallucination detectors suffer from the low-resource effect? We conduct experiments on five tasks across three domains (factual recall, STEM, and Humanities). Experiments with four LLMs and three hallucination detectors reveal a curious finding: As expected, the task accuracies in low-resource languages experience large drops (compared to English). However, the drop in detectors' accuracy is often several times smaller than the drop in task accuracy. Our findings suggest that even in low-resource languages, the internal mechanisms of LLMs might encode signals about their uncertainty. Further, the detectors are robust within language (even for non-English) and in multilingual setups, but not in cross-lingual settings without in-language supervision.

幻觉检测低资源语言多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。