arXiv:2606.11105cs.CLcs.AI2026-06被引 1

测试大模型对虚构概念的识别能力,发现多数模型竟会信口开河。

PhantomBench: Benchmarking the Non-existential Threat of Language Models

论文配图:PhantomBench: Benchmarking the Non-existential Threat of Language Models
图 1 · 摘自论文原文
  • 构建6万+虚构术语库,模拟真实但不存在的概念
  • 21个模型平均幻觉率高达86.7%,顶尖模型也难免疫
  • 可用来评估模型对冷门概念的泛化能力,适合安全研究者

幻觉问题——语言模型生成与事实不符的内容——带来严重风险,因用户常盲目信赖其输出。尤其在高风险领域,此类行为可能导致重大危害。尽管对幻觉现象已有一定研究,但模型识别自身知识边界的能力仍不清晰。我们提出PhantomBench,首个大规模基准,包含超过6万条来自不同领域的虚构术语和实体。通过该基准,我们评估了21种不同类型和规模的模型。结果显示,所有模型幻觉率惊人(部分高达86.7%),甚至前沿模型在输入暗示虚构概念存在时也不愿拒绝回答。我们进一步证明,PhantomBench可作为研究模型对罕见概念幻觉行为的代理工具。此外,我们提供可扩展的构建流程,支持研究人员按需生成定制化虚构概念。

原文摘要 · Abstract (English)

Hallucinations, where language models (LMs) generate factually ungrounded responses, pose serious risks, as users tend to blindly rely on them. This is particularly concerning in high-stakes domains, where consequences of such model behavior can lead to significant harms. Despite notable progress in understanding hallucinations, it remains unclear how reliably these models can recognize the limits of their knowledge. We introduce PhantomBench, the first large-scale benchmark of its kind, comprising more than 60K non-existent terms and entities derived from real concepts across diverse domains. Using our benchmark, we evaluate a total of 21 models of various types and sizes. We show staggering hallucination rates across the board (with average rates as high as 86.7% in some cases), and note that even frontier models surprisingly fail to abstain on non-existent concepts, especially when the input presumes their existence. We then show that PhantomBench can serve as a proxy for studying model behavior on rare concepts for which models are more prone to hallucinate. We also provide a pipeline to construct PhantomBench, enabling scalable generation of non-existent concepts tailored to the specific needs of researchers and practitioners.

幻觉检测语言模型评测基准虚构概念

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。