揭示视觉语言模型的'视而不见'问题,提出量化'看见代价'的新评估方法。
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm

- 用语义迁移协议替代数据删减,定量评估视觉信息利用程度
- 提出三类新指标:看见税、看见诅咒、看见谬误,统一衡量多模态可信度
- 发现大语言模型越强,视觉瓶颈惩罚可能越大,警示盲目扩张风险
视觉语言模型(VLMs)的快速普及常被视作实现统一多模态知识发现的途径,但其背后隐含一个未经检验的假设:现有VLM能忠实融合多模态数据。我们指出,这些模型往往无法做到,反映出主流视觉编码器-投影器-大语言模型范式中的可信度问题。顶尖模型常表现出功能盲区,即依赖强大的语言先验绕过严重的视觉表征瓶颈,而非从视觉输入中提取真实知识。本文挑战了传统多模态评估方法——依赖数据删减或新数据集创建,从而混淆了数据偏见与架构能力不足。我们提出一种信息论框架:模态翻译协议,用于量化‘看见代价’。通过迁移语义负载而非删减,我们定义了三项新指标:看见税(ToS)、看见诅咒(CoS)、看见谬误(FoS),最终形成语义充分性准则(SSC)。此外,我们提出多模态扩展的发散定律:当底层语言引擎规模扩大至前所未有的推理能力时,视觉知识瓶颈的代价反而可能上升。我们主张,社区应超越‘多模态增益’作为主要评估目标。将SSC从被动诊断约束转变为积极架构蓝图,为下一代真正具备多模态推理能力的AI系统奠定基础。
原文摘要 · Abstract (English)
The rapid proliferation of Vision-Language Models (VLMs) is often framed as enabling unified multimodal knowledge discovery but rests on an under-examined assumption: that current VLMs faithfully synthesise multimodal data. We argue they often do not, and this gap reflects a trustworthiness problem in the dominant Vision Encoder-Projector-LLM paradigm. Rather than extracting grounded knowledge from visual inputs, state-of-the-art models frequently exhibit functional blindness, i.e., exploiting strong language priors to bypass severe visual representation bottlenecks. In this work, we challenge the conventional methodology of multimodal evaluation, which relies on data ablation or new dataset creation and therefore conflates dataset biases with architectural incapacity. We propose an information-theoretic departure: the Modality Translation Protocol, designed to quantify what we call the Expense of Seeing. By translating semantic payloads rather than ablating them, we formulate three novel metrics -- the Toll (ToS), Curse (CoS), and Fallacy (FoS) of Seeing -- culminating in the Semantic Sufficiency Criterion (SSC). Furthermore, we hypothesise a Divergence Law of Multimodal Scaling: as the underlying language engines scale to unprecedented reasoning capabilities, the penalty of the visual knowledge bottleneck may increase rather than diminish. We argue the community should move beyond "multimodal gain" as a primary evaluation target. By elevating the SSC from a passive diagnostic constraint to an active architectural blueprint, we provide a foundation for guiding the next generation of AI systems toward genuine multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。