arXiv:2601.15511cs.CLcs.CY2026-01被引 1

首个针对高风险领域对抗性事实性的评测基准,检验大模型识破伪装成可信谎言的能力。

AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains

  • 构建健康、金融、法律三领域对抗性事实性测试集,分难易两档评估模型防御力。
  • 80B版本的Qwen3在去除非意义回复后准确率最高,GPT-5表现稳定且整体领先。
  • 发现模型规模与性能非线性增长,长文本输出受注入错误信息影响不显著。

大型语言模型(LLMs)中的幻觉问题仍属严峻挑战,尤其在高风险领域加剧了错误信息传播并削弱公众信任。其中,事实性幻觉关乎模型与已知世界知识的一致性。对抗性事实性指在提示中以不同置信度刻意插入虚假信息,测试模型识别和抵抗此类高信心谎言的能力。现有研究缺乏高质量、领域特定的评估资源,且未考察注入错误信息对长文本事实性的影响。为此,我们提出AdversaRiskQA,首个经验证的、可靠的跨健康、金融、法律领域的对抗性事实性基准。该基准包含两个难度层级,用于评估模型在不同知识深度下的防御能力。我们提出两种自动化方法评估攻击成功率与长文本事实性。评测了来自Qwen、GPT-OSS和GPT系列的六款开源与闭源模型,测量其误判检测率。长文本事实性在Qwen3(30B)上进行基线与对抗条件下的评估。结果显示,在排除无意义回复后,Qwen3(80B)平均准确率最高,而GPT-5保持持续高精度。性能随模型规模呈非线性增长,领域间差异明显,难度差距随模型增大而缩小。长文本评估显示,注入的错误信息与模型输出事实性之间无显著相关性。AdversaRiskQA为定位大模型弱点、开发更可靠高风险应用模型提供了重要工具。

原文摘要 · Abstract (English)

Hallucination in large language models (LLMs) remains an acute concern, contributing to the spread of misinformation and diminished public trust, particularly in high-risk domains. Among hallucination types, factuality is crucial, as it concerns a model's alignment with established world knowledge. Adversarial factuality, defined as the deliberate insertion of misinformation into prompts with varying levels of expressed confidence, tests a model's ability to detect and resist confidently framed falsehoods. Existing work lacks high-quality, domain-specific resources for assessing model robustness under such adversarial conditions, and no prior research has examined the impact of injected misinformation on long-form text factuality. To address this gap, we introduce AdversaRiskQA, the first verified and reliable benchmark systematically evaluating adversarial factuality across Health, Finance, and Law. The benchmark includes two difficulty levels to test LLMs' defensive capabilities across varying knowledge depths. We propose two automated methods for evaluating the adversarial attack success and long-form factuality. We evaluate six open- and closed-source LLMs from the Qwen, GPT-OSS, and GPT families, measuring misinformation detection rates. Long-form factuality is assessed on Qwen3 (30B) under both baseline and adversarial conditions. Results show that after excluding meaningless responses, Qwen3 (80B) achieves the highest average accuracy, while GPT-5 maintains consistently high accuracy. Performance scales non-linearly with model size, varies by domains, and gaps between difficulty levels narrow as models grow. Long-form evaluation reveals no significant correlation between injected misinformation and the model's factual output. AdversaRiskQA provides a valuable benchmark for pinpointing LLM weaknesses and developing more reliable models for high-stakes applications.

对抗性评测事实性高风险领域大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。