发现模型答对题时内部状态仍可能紧张,揭示了评估盲区。
When Correct Decisions Hide Internal Stress: Decision-State Probing in Multimodal Language Models

- 设计新评测框架S³E,通过语义压力对比测试内部决策稳定性。
- 即使选择正确,模型在关键层仍出现显著状态偏移,说明内部不稳。
- 适合关注模型可信度与内部机制的研究者,尤其多模态领域。
多模态语言模型通常通过外部行为评估:正确匹配图文、拒绝无依据描述或准确回答视觉问题。然而,正确行为并不意味着模型内部决策状态在语义压力下保持稳定。我们通过S³E(结构化语义压力评估)框架研究这一差距。S³E采用正锚定A/B强制选择设置,将图像支持的描述与语义压力候选项对比,同时在原始和交换选项顺序下进行测试,并提取预答案决策阶段的隐藏状态。聚焦于严格正确的试验(即模型在两种顺序下均选择正确描述),我们不将任意隐藏状态变化视为不稳,而是衡量语义冲突候选是否引起相对于语义保留对照组的超额决策状态位移。在Qwen3VL、Gemma3和InternVL3上,语义压力始终导致选定层的超额位移,而随机负样本对比结果则因模型而异。我们将其解释为有限范围内的决策状态压力敏感信号,而非下游失败或幻觉证据。结果表明,强制选择正确性本身不足以证明内部决策几何不变。
原文摘要 · Abstract (English)
Multimodal language models are typically evaluated through external behavior: selecting the correct image--text match, rejecting unsupported captions, or answering visual queries correctly. However, correct behavior alone does not show that the model's internal decision state remains stable under controlled semantic stress. We study this gap through S$^3$E (Structured Semantic Stress Evaluation), a framework for analyzing behavior-internal decoupling in multimodal language models. S$^3$E uses a positive-anchored A/B forced-choice setup in which an image-supported caption is contrasted against semantic stress candidates under both original and swapped option orders, while hidden states are extracted at the pre-answer decision state. We focus on strict-correct trials, where the model consistently selects the correct caption across both orders. Rather than treating arbitrary hidden-state variation as evidence of instability, we measure whether semantic-conflict candidates induce excess decision-state displacement relative to meaning-preserving controls. Across Qwen3VL, Gemma3, and InternVL3, semantic stress consistently produces positive selected-layer excess displacement over lexical controls despite correct forced-choice behavior, while comparisons against random negatives are model-dependent. We interpret this as a scoped decision-state stress-sensitivity signal rather than evidence of downstream failure or hallucination. Our results suggest that forced-choice correctness alone is not a sufficient certificate of invariant internal decision geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。