发现多模态大模型感知对了却仍会错误回应,问题出在如何把感知结果转成正确行动。
Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

- 设计新评测集IMAVB,分离测试模型识别文本与感官输入冲突的能力
- 8个模型均存在感知正确但回答错误的‘表征-行为断层’现象
- 音频模态比视觉更难准确接地,且提示工程无法有效改善该缺陷
当多模态大语言模型面对一个与自身所见或所听相矛盾的文本前提时,失败根源在于感知还是行为?尽管近期多模态模型被视作基于感知的智能体,能联合处理视频、音频和文本,但一项基本的感知对齐能力仍未被检验:即识别文本陈述与模型自身感官输入之间的冲突。本文提出IMAVB,一个由500段长视频组成的精选基准数据集,采用2×2实验设计,交叉考察目标模态(视觉、音频)与前提条件(标准、误导性),从而可独立衡量冲突检测能力与常规多模态理解能力。在8个开源多模态大模型及Gemini 3.1 Pro上测试发现存在显著的‘表征-行为断层’:隐藏状态能可靠编码前提与感知不一致信号,但模型几乎从不拒绝错误主张。行为层面,模型呈现两种失效模式:低估拒绝(under-rejection),将误导性问题当作真命题回应;高估拒绝(over-rejection),频繁拒绝标准问题,牺牲正常理解准确性。该断层具有模态不对称性(音频接地表现弱于视觉),且在七种提示变体下均难以缓解。作为初步诊断干预,一种探针引导的逻辑值调整(PGLA)方法重新注入编码的不一致信号至解码过程,显著提升拒绝行为。结果表明,多模态对齐的瓶颈在于信息翻译,而非感知本身。
原文摘要 · Abstract (English)
When an omnimodal large language model accepts a question whose textual premise contradicts what it actually sees or hears, does the failure lie in perception or in action? Recent omnimodal models are positioned as perception-grounded agents that jointly process video, audio, and text, yet a basic form of grounding remains untested: catching a textual claim that conflicts with the model's own sensory input. We introduce IMAVB, a curated 500-clip benchmark of long-form movies with a 2x2 design crossing target modality (vision, audio) and premise condition (standard, misleading), which lets us measure conflict detection separately from ordinary multimodal comprehension. Across eight open-source omnimodal LLMs and Gemini 3.1 Pro, we document a Representation-Action Gap: hidden states reliably encode premise-perception mismatches even when the same models almost never reject the false claim in their outputs. Behaviorally, models fall into two failure modes: under-rejection, in which they answer misleading questions as if the false premise were true; and over-rejection, in which they reject more often but also reject standard questions, sacrificing ordinary comprehension accuracy. The gap is modality-asymmetric (audio grounding underperforms vision) and prompt-resistant across seven variants. As an initial diagnostic intervention, a probe-guided logit adjustment (PGLA) re-injects the encoded mismatch signal into decoding and consistently improves rejection behavior. Together, these results suggest the bottleneck for omnimodal grounding lies in translation, not perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。