arXiv:2507.23453cs.CRcs.CL2025-07被引 1

提出反事实评估框架,提升大模型评估系统对盲攻击的检测能力。

Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems

  • 在标准评估外引入反事实评估,用假答案重新验证提交结果
  • 在标准和反事实条件下均通过验证的攻击可被有效识别
  • 大幅增强安全性,且对正常评估性能影响极小

本文研究大模型评估系统对提示注入攻击的防御方法。我们定义了一类称为盲攻击的威胁:候选答案独立于真实答案生成,以欺骗评估器。为应对此类攻击,提出一种框架,在标准评估(SE)基础上增加反事实评估(CFE),即用故意错误的真值答案重新评估提交内容。若系统在标准与反事实条件下均认可同一答案,则判定为攻击。实验表明,标准评估极易被攻破,而所提的SE+CFE框架显著提升安全性,实现高检测率且性能损失微乎其微。

原文摘要 · Abstract (English)

This paper investigates defenses for LLM-based evaluation systems against prompt injection. We formalize a class of threats called blind attacks, where a candidate answer is crafted independently of the true answer to deceive the evaluator. To counter such attacks, we propose a framework that augments Standard Evaluation (SE) with Counterfactual Evaluation (CFE), which re-evaluates the submission against a deliberately false ground-truth answer. An attack is detected if the system validates an answer under both standard and counterfactual conditions. Experiments show that while standard evaluation is highly vulnerable, our SE+CFE framework significantly improves security by boosting attack detection with minimal performance trade-offs.

大模型安全评估防御反事实评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。