arXiv:2608.06571cs.CRcs.CL2026-08

研究视觉语言模型在答案不变攻击下的置信度可靠性,发现其易被操纵且无法提供有效监督。

Model Confidence Under Answer-Preserving Attacks: An Informativeness-Manipulability Frontier

  • 在保持答案字节一致的前提下,通过对抗扰动破坏置信度读出。
  • 84种组合中,所有置信度通道均被攻破,最低判别率降至0.617的基线水平。
  • 适合关注AI安全、模型可信度与防御评估的研究者阅读。

部署的视觉语言系统常以置信度作为答案输出的门控机制,因此置信度鲁棒性对监督至关重要。本文研究在白盒、仅图像攻击下,置信度读出在保持生成答案字节完全一致的约束条件下的表现。在可达性假设下,不可移动的读出性能无法优于答案字符串准确率基线(0.617)。独立于该假设,若存在低于可测阈值的均匀振幅证书,则可保证对抗判别性能不低于该基线。在四个视觉问答模型、三个基准数据集、五个部署置信通道及两个防御估计器上,直接或代理目标攻击均产生可行的逐项扰动,在全部84个估计器-单元组合中推翻该均匀证书。协同的正确性标签感知攻击使对抗判别性能降至或低于答案字符串基线,涵盖所有60个部署通道单元,包括59个初始高于基线的单元。隐藏状态干预与开放式文本模型激活空间复现表明,类似置信度变化可在表示层而非仅通过对抗图像实现。四种测试防御家族均未在此评估设置下建立稳健替代方案。在置信度门控模拟中,协同的词概率攻击迁移至隐藏状态门控后,使原本被拒绝的错误答案中有高达84.8%被接受。重加权至各基准自然正确率后,转移攻击下12个单元中有8个接受准确率低于无门控基线,直接门控攻击下全部12个单元均低于基线。在所研究威胁模型与预算下,置信度是敏感于完整性而非内在鲁棒的监督信号。

原文摘要 · Abstract (English)

Deployed vision-language systems often gate their answers on confidence, making confidence robustness relevant to oversight. We study confidence readouts under white-box, image-only attacks constrained to preserve the generated answer byte-identically. Under a reachability assumption, an unmovable readout cannot outperform the answer-string accuracy prior, whose pooled value is 0.617. Independently of that assumption, a uniform amplitude certificate below a measurable threshold guarantees adversarial discrimination above the same floor. Across four vision-language models, three visual question answering benchmarks, five deployed confidence channels and two defense estimators, direct or surrogate-aimed attacks produce itemwise feasible perturbations that refute this uniform certificate in all 84 estimator-by-cell combinations. Coordinated correctness-label-aware attacks drive adversarial discrimination to or below the answer-string floor in all sixty deployed-channel cells, including all fifty-nine that begin above it. Hidden-state interventions and an open-ended text-model activation-space replication show that comparable confidence movement can be induced at the representation level rather than only through adversarial images. None of four tested defense families establishes a robust alternative under the specific evaluation applied to it. In a confidence-gated simulation, a coordinated token-probability attack transferred to a hidden-state gate causes up to 84.8% of previously rejected wrong answers to become accepted. After reweighting to each benchmark's natural correctness prevalence, accepted accuracy falls below the no-gate baseline in eight of twelve cells under transfer and all twelve under a direct gate-aimed attack. Under the studied threat model and budget, confidence is therefore an integrity-sensitive rather than intrinsically robust oversight signal.

模型安全置信度对抗攻击视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。