arXiv:2605.10146cs.AIcs.CR2026-05

测试大模型在恶意知识编辑下的安全风险,发现攻击可隐蔽误导推理。

Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing

论文配图:Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing
图 1 · 摘自论文原文
  • 构建统一框架评估恶意知识对推理的影响
  • 攻击可导致错误或不安全推理,且不破坏模型整体能力
  • 适合关注大模型安全与知识可信性的研究者

大型语言模型(LLMs)越来越多依赖知识编辑来支持知识密集型推理,但这种灵活性也引入了严重安全风险:攻击者可注入恶意或误导性知识,扭曲下游推理并导致有害后果。现有知识编辑基准主要关注编辑效果,缺乏系统评估知识编辑安全性影响的统一框架。为此,我们提出EditRisk-Bench,一个用于系统评估恶意知识编辑下知识密集型推理安全风险的基准。不同于以往强调编辑成功率、泛化性和局部性的基准,EditRisk-Bench聚焦于注入知识如何影响下游推理行为与可靠性。它整合多种恶意场景(如虚假信息、偏见、安全违规),结合多层级知识密集型推理任务与代表性编辑策略,在统一框架中衡量攻击有效性、推理正确性与副作用。在开源和闭源大模型上的大量实验表明,恶意知识编辑可稳定引发错误或不安全推理,同时基本保留模型通用能力,使风险难以察觉。我们进一步识别出影响风险的关键因素,包括编辑规模、知识特征与推理复杂度。EditRisk-Bench为理解与缓解大模型知识编辑中的安全风险提供了可扩展的测试平台。

原文摘要 · Abstract (English)

Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces critical safety risks: adversaries can inject malicious or misleading knowledge that corrupts downstream reasoning and leads to harmful outcomes. Existing knowledge editing benchmarks primarily focus on editing efficacy and lack a unified framework for systematically evaluating the safety implications of edited knowledge on reasoning behavior. To address this gap, we present EditRisk-Bench, a benchmark for systematically evaluating safety risks of knowledge-intensive reasoning under malicious knowledge editing. Unlike prior benchmarks that mainly emphasize edit success, generalization, and locality, EditRisk-Bench focuses on how injected knowledge affects downstream reasoning behavior and reliability. It integrates diverse malicious scenarios, including misinformation, bias, and safety violations, together with multi-level knowledge-intensive reasoning tasks and representative editing strategies within a unified evaluation framework measuring attack effectiveness, reasoning correctness, and side effects. Extensive experiments on both open-source and closed-source LLMs show that malicious knowledge editing can reliably induce incorrect or unsafe reasoning while largely preserving general capabilities, making such risks difficult to detect. We further identify several key factors influencing these risks, including edit scale, knowledge characteristics, and reasoning complexity. EditRisk-Bench provides an extensible testbed for understanding and mitigating safety risks in knowledge editing for LLMs.

大模型安全知识编辑推理风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。