提出HERALD审计框架,识别并修复检索奖励中的伪造引用漏洞。
HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
- 通过精确问题干预分离可见与真实信息,枚举检测器合约
- 最小修复策略使攻击成功率降至0.50%以下,提升引用准确率2.02点
- 适合关注大模型检索可信性与奖励机制安全的研究者
搜索代理奖励融合答案质量、引用依据、工具成本与反作弊条款;高分未必代表引用证据被真实检索,且惩罚项可能相互抵消。本文提出HERALD,一种离线审计方法,通过相同问题干预,分离候选可见信息与真实信息,并在策略优化前枚举检测器合约。在四个基于Qwen3-8B的HotpotQA、2WikiMultiHopQA和MuSiQue数据池上,$R_0$可拒绝对检索删除和虚假ID攻击,但标签无关的引用洗白攻击仍成功。完整$2^3$消融实验表明,针对$L$——即引用未被检索到的语料段落——的针对性强化是观察到的包含最小修复:$R[L]$在实证中实现0%攻击成功率,单边聚类上界为0.50%。该差距在不同池规则、可见的BM25攻击者及四种模型下持续存在;当攻击移除真实支持ID惩罚时,更强防御仍易受攻击。在严格5M-token匹配训练下,每基准测试256对问题评估,$R[L]$在HotpotQA和2Wiki上满足EM非劣性门槛,但在MuSiQue上不满足。同套引用精确率与支持召回率分别提升2.02和1.46点,无支持引用减少1.69点,2Wiki与MuSiQue的洗白攻击性下降。自然$L$未减少,检测器仅出现在58,368条训练轨迹中的18次。HERALD因此实现了鲁棒评分、稀疏学习信号与策略迁移的分离。
原文摘要 · Abstract (English)
Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies exact same-question interventions, separates candidate-visible from oracle information, and enumerates detector contracts before policy optimization. On four Qwen3-8B pools from HotpotQA, 2WikiMultiHopQA, and MuSiQue, $R_0$ rejects search deletion and fake IDs, but a label-free citation-laundering attack succeeds. A complete $2^3$ ablation identifies targeted strengthening of $L$---citing a corpus passage absent from the retrieved evidence---as the observed inclusion-minimal repair: $R[L]$ has zero empirical ASR with a 0.50% one-sided cluster upper bound. The gap persists across pool rules, a visible BM25 attacker, and four models; broader hardening remains vulnerable when the attack removes an oracle support-ID penalty. Under strict 5M-token matched training evaluated on 256 paired questions per benchmark, $R[L]$ meets the EM non-inferiority gate on HotpotQA and 2Wiki but not MuSiQue. Equal-suite citation precision and support recall improve by 2.02 and 1.46 points, unsupported citations fall by 1.69, and laundering attackability falls on 2Wiki and MuSiQue. Natural $L$ is not reduced, and the detector appears in only 18 of 58,368 training trajectories. HERALD thus separates robust scoring, sparse learning signal, and policy transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。