提出新评估方法,精准检测伪造声明改写是否真能骗过验证器。
AtomEval: Validity-Aware Atomic Evaluation of Adversarial Claim Rewriting in Fact Verification
- 将声明拆解为原子单元,区分真假改写与有效欺骗。
- 实测发现传统成功率虚高,部分改写反而修正了原错误命题。
- 适合研究对抗性改写、模型安全性的研究人员使用。
大语言模型可重写被证伪的声明以逃避基于证据的事实验证,但传统攻击成功率(ASR)在改写改变、弱化或修正原错误命题时会被夸大。我们提出AtomEval,一种针对固定证据的对抗性声明重写的有效性感知评估协议。AtomEval将声明表示为实体-关系-对象-修饰符(SROM)原子,通过单向保留门分离出真正规避验证器的改写与改变命题的改写,并报告有效性感知攻击成功率(VASR),仅统计既规避验证又保持原始错误命题的改写。该方法还提供细粒度诊断,解释命题层面的失败和非最小有效改写。在FEVER数据集上的实验证明,许多看似成功的攻击实际改变了原命题;通过显式且可度量地要求保留被攻击命题,AtomEval为对抗性重写器提供了稳定评估目标,使其需在规避验证与保持命题不变之间取得平衡。
原文摘要 · Abstract (English)
Large language models (LLMs) can rewrite refuted claims to evade evidence-based fact verifiers, but conventional attack success rate (ASR) can be inflated when rewrites change, weaken, or correct the false proposition they are supposed to preserve. We introduce AtomEval, a validity-aware evaluation protocol for fixed-evidence adversarial claim rewriting. AtomEval represents claims as subject--relation--object--modifier (SROM) atoms, applies a one-way preservation gate to separate valid verifier evasion from proposition-changing rewrites, and reports validity-aware attack success rate (VASR), which counts only verifier-evasive rewrites that preserve the original false proposition. AtomEval further provides fine-grained diagnostics that explain both proposition-level failures and non-minimal valid rewrites. On FEVER refuted-claim rewriting, AtomEval exposes and explains ASR inflation: many apparent attacks fool the verifier by altering, weakening, or correcting the proposition they should preserve. By making attacked-proposition preservation explicit and measurable, AtomEval provides a stable evaluation target for evaluating adversarial rewriters that must balance verifier evasion with proposition preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。