arXiv:2510.01782cs.CLcs.AI2025-10中稿 · ICLR被引 6

提出新指标衡量大模型对未知问题的拒绝能力,更真实反映其事实可靠性。

Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks

  • 用斯皮尔曼相关性定义拒绝指数,评估模型拒答与错误的匹配度。
  • 16个模型5个数据集测试显示,该指标稳定且不依赖准确率或拒答率。
  • 揭示大模型虽答得准,但拒答行为可能不可靠,适合关注可信性的研究者。

大型语言模型(LLMs)应拒绝回答超出自身知识范围的问题,这种能力我们称为知识感知拒绝,对事实可靠性至关重要,但现有度量方法无法有效捕捉。本文提出拒绝指数(RI),一种新颖且严谨的度量标准,用于衡量模型对未知问题拒绝的准确性。RI定义为拒绝概率与错误概率之间的斯皮尔曼等级相关系数。通过轻量级双轮评估法,仅需两次标准评估中的观察拒答率即可实际测量。在16个模型和5个数据集上的广泛实验表明,RI能准确量化模型的知识感知拒绝能力。值得注意的是,RI在不同拒答率下保持稳定,且模型排名不受整体准确率和拒答率影响。这些特性表明,RI捕捉到了模型知识校准的一个稳定内在属性。更重要的是,RI揭示了大模型事实性中一个此前被忽视的重要方面:尽管大模型在事实任务上表现高准确率,其拒绝行为却可能不可靠且脆弱。

原文摘要 · Abstract (English)

Large Language Models (LLMs) should refuse to answer questions beyond their knowledge. This capability, which we term knowledge-aware refusal, is crucial for factual reliability, while existing metrics fail to capture this ability. In this work, we propose the Refusal Index (RI), a novel and principled metric that measures how accurately LLMs refuse questions they do not know. We define RI as Spearman's rank correlation between refusal probability and error probability. RI is practically measurable with a lightweight two-pass evaluation method which only require observed refusal rates across two standard evaluation runs. Extensive experiments across 16 models and 5 datasets demonstrate that RI accurately quantifies a model's knowledge-aware refusal capability. Notably, RI remains stable across different refusal rates and provides consistent model rankings independent of a model's overall accuracy and refusal rates. These properties suggest RI captures a stable, intrinsic aspect of model knowledge calibration. More importantly, RI provides insight into an important but previously overlooked aspect of LLM factuality: while LLMs achieve high accuracy on factual tasks, their refusal behavior can be unreliable and fragile.

大模型拒绝机制知识校准评估指标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。