arXiv:2511.09971cs.CL2025-11The 14th Internati…

通过数值扰动测试大模型在真假判断中的脆弱性,发现主流模型准确率最高下降62%。

NumPert: Numerical Perturbations to Probe Language Models for Veracity Prediction

  • 用可控的数值扰动评估语言模型对事实判断任务的鲁棒性
  • 顶级模型在特定扰动下准确率下降最高达62%,无一完全稳健
  • 增加上下文长度会降低准确率,但加入扰动示范可显著恢复性能

大型语言模型在知识密集型任务(如事实核查、问答)中表现优异,但在数值推理方面仍存在困难。本文通过受控扰动(包括标签翻转探测)系统评估了当前最先进的模型在数值声明与证据配对上的真假判断能力。结果显示,即使领先专有系统在特定扰动下准确率最高下降62%。没有任何模型能在所有条件下保持稳健。进一步发现,增加上下文长度通常会降低准确率,但若扩展上下文包含扰动示范,则多数模型能显著恢复性能。这些结果揭示了当前数值事实核查中的关键局限,表明鲁棒性仍是语言模型面临的重要挑战。

原文摘要 · Abstract (English)

Large language models show strong performance on knowledge intensive tasks such as fact-checking and question answering, yet they often struggle with numerical reasoning. We present a systematic evaluation of state-of-the-art models for veracity prediction on numerical claims and evidence pairs using controlled perturbations, including label-flipping probes, to test robustness. Our results indicate that even leading proprietary systems experience accuracy drops of up to 62\% under certain perturbations. No model proves to be robust across all conditions. We further find that increasing context length generally reduces accuracy, but when extended context is enriched with perturbed demonstrations, most models substantially recover. These findings highlight critical limitations in numerical fact-checking and suggest that robustness remains an open challenge for current language models.

语言模型事实核查数值推理鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。