伪造的同行评价可误导多模态大模型评审团,新方法能有效识别并防御此类攻击。
Forged Peer Judgments Mislead Multimodal LLM Judge Panels: Source-Blind Anchoring and Panel-Consensus Verification

- 通过伪造简洁错误引用,可大幅扭曲评审结果
- 真实错误引用被伪造引用影响的几率高1.5至2.7倍
- 引入盲评共识验证机制可拦截84.9%的伪造攻击
多模态大模型评审团虽可交叉参考同行意见,但被引用的判断本身可能不可信。我们揭示了视觉-语言模型(VLM)评审团中存在文本层面的‘源盲锚定’攻击面。在自评与同行评述两种情境下,引用独立视觉判断会产生19至26个百分点的锚定偏差。匹配内容仅标签不同的对照实验显示,错误率仅变化-0.17个百分点(95%置信区间[-0.68, 0.35]),说明偏差并非由自评/同行标签导致。在测试构造中,人为生成的简短错误引用,比自然发生的错误同行陈述更易推翻原本正确的裁决,影响概率高出1.5至2.7倍,且双数据集、七种VLM裁判下置信区间均不包含等效性。因两类陈述在选择和形式上存在差异,该比率反映的是特定攻击下的破坏程度,而非单纯溯源因果。随后提出面板共识验证,通过独立盲评交叉核验引用内容。该方法可阻断84.9%的伪造攻击,将净伤害降低97.5%,并在留一法重验证下仍保留真实同行信息的正向但统计上不显著的点估计。结果揭示了一个低成本攻击面及切实可行的防御方案,提升多模态协同评估的安全性。
原文摘要 · Abstract (English)
Multimodal LLM judge panels can cross-reference peers, but a quoted peer judgment may itself be untrusted. We expose source-blind anchoring as a text-level attack surface in vision-language model (VLM) panels. Quoting independent visual judgments creates large anchoring gaps (19--26 percentage points) under both self and peer framing. A matched-content, label-only control changes the broken rate by only $-0.17$pp (95\% CI $[-0.68,0.35]$), showing that the self/peer label itself does not explain the effect. Under our tested construction, deliberately generated, concise wrong quotes overturn originally-correct verdicts 1.5--2.7$\times$ more often than naturally occurring wrong peer statements, with bootstrap 95\% CIs excluding parity across two datasets and seven VLM judges. Because the two statement populations differ in selection and form, this ratio measures differential damage under the tested attack rather than a provenance-only causal effect. We then introduce panel-consensus verification, which cross-checks a quote against independently collected blind votes. It blocks 84.9\% of fabricated attacks, cuts their net harm by 97.5\%, and preserves the positive but statistically inconclusive point estimate for genuine peer information under leave-one-out re-verification. These results identify a low-cost attack surface and a concrete defense for safer multimodal collaborative evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。