用压力信号对抗视觉语言模型的错误承诺,提升回答可靠性。
TPCD: Tone-Pressure Contrastive Decoding and the Label-Free Gating Bottleneck in Vision-Language Models

- 通过对比高压与安全指令下的输出差异,动态修正模型响应
- 在800样本测试中将攻击成功率从66.75%降至1.63%
- 适合关注模型鲁棒性与安全性的研究人员使用
高压提示会诱使视觉语言模型(VLMs)产生不当承诺,如识别模糊文字、误报时间或确认不存在物体。本文提出音调-压力对比解码(TPCD),通过减去高压指令下的logits与安全中性指令下的logits,实现对抗性修正。在800例的音调敏感基准上,LLaVA-1.5-7B在高压下攻击成功率达66.75%;安全中性化将其降至9.88%;完整TPCD虽将攻击率压至0.50%,但正面准确率跌至15.56%。基于任务优先/分歧的门控机制,在保持正面准确率54.44%的同时,将攻击率降至1.63%。在GLM-4.6V和Llama-3.2-Vision上进行独立验证,简单门控优于安全中性化,敏感性分析限定弱时间正向子任务边界。无类别先验的答案分歧路由进一步将整体攻击率降至6.93%,优于安全中性化(10.98%)与分支分歧(9.67%),且维持79.94%正面准确率,但仍是事后、表层特征依赖的方法。结论:压力是检测承诺偏见的有效探针,也是可行的缓解信号,但现有门控尚未被独立验证为具备语义感知的检测器。
原文摘要 · Abstract (English)
High-pressure prompts can push vision-language models (VLMs) into unsupported commitments, such as reading illegible text, reporting indeterminate times, or affirming absent objects. This paper asks whether the pressure-induced distribution itself can serve as a contrastive-decoding negative branch. Tone-pressure contrastive decoding (TPCD) subtracts logits produced under a high-pressure instruction from logits produced under a safe neutral instruction. On the 800-example tone-matters benchmark, LLaVA-1.5-7B under pressure reaches 66.75% attack success rate (ASR); safe neutralization reduces ASR to 9.88%; full TPCD reaches 0.50% but collapses positives to 15.56%. A benchmark-specific task-prior/disagreement gate preserves measured positive accuracy (54.44%) while lowering ASR to 1.63% on LLaVA. Treating this LLaVA analysis as the design split, full $n=800$ negative and $n=780$ matched-positive held-out runs on GLM-4.6V and Llama-3.2-Vision show that simple gates can improve over safe neutralization, with sensitivity analyses bounding the weak time-positive subtask. A category-prior-free answer-disagreement router reduces held-out aggregate ASR to 6.93%, improving over both safe neutralization (10.98%) and branch disagreement (9.67%) while matching branch disagreement's 79.94% positive accuracy, although it remains post-hoc and surface-form based. We conclude that pressure is a useful probe of commitment bias and a viable mitigation signal, but the current gates are not yet independently validated grounding-aware detectors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。