arXiv:2604.07967cs.CLcs.AI2026-04

提出新评估方法,精准检测伪造声明改写是否真能骗过验证器。

AtomEval: Validity-Aware Atomic Evaluation of Adversarial Claim Rewriting in Fact Verification

  • 将声明拆解为原子单元,区分真假改写与有效欺骗。
  • 实测发现传统成功率虚高,部分改写反而修正了原错误命题。
  • 适合研究对抗性改写、模型安全性的研究人员使用。

大语言模型可重写被证伪的声明以逃避基于证据的事实验证,但传统攻击成功率(ASR)在改写改变、弱化或修正原错误命题时会被夸大。我们提出AtomEval,一种针对固定证据的对抗性声明重写的有效性感知评估协议。AtomEval将声明表示为实体-关系-对象-修饰符(SROM)原子,通过单向保留门分离出真正规避验证器的改写与改变命题的改写,并报告有效性感知攻击成功率(VASR),仅统计既规避验证又保持原始错误命题的改写。该方法还提供细粒度诊断,解释命题层面的失败和非最小有效改写。在FEVER数据集上的实验证明,许多看似成功的攻击实际改变了原命题;通过显式且可度量地要求保留被攻击命题,AtomEval为对抗性重写器提供了稳定评估目标,使其需在规避验证与保持命题不变之间取得平衡。

原文摘要 · Abstract (English)

Large language models (LLMs) can rewrite refuted claims to evade evidence-based fact verifiers, but conventional attack success rate (ASR) can be inflated when rewrites change, weaken, or correct the false proposition they are supposed to preserve. We introduce AtomEval, a validity-aware evaluation protocol for fixed-evidence adversarial claim rewriting. AtomEval represents claims as subject--relation--object--modifier (SROM) atoms, applies a one-way preservation gate to separate valid verifier evasion from proposition-changing rewrites, and reports validity-aware attack success rate (VASR), which counts only verifier-evasive rewrites that preserve the original false proposition. AtomEval further provides fine-grained diagnostics that explain both proposition-level failures and non-minimal valid rewrites. On FEVER refuted-claim rewriting, AtomEval exposes and explains ASR inflation: many apparent attacks fool the verifier by altering, weakening, or correcting the proposition they should preserve. By making attacked-proposition preservation explicit and measurable, AtomEval provides a stable evaluation target for evaluating adversarial rewriters that must balance verifier evasion with proposition preservation.

对抗攻击事实验证评估方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。