arXiv:2606.01679cs.CL2026-06

模型能看懂图表数据却用不上,关键在信息路由失败。

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

论文配图:Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
图 1 · 摘自论文原文
  • 分析模型中间层发现图表信息被编码但未传到预测位置。
  • 同一数据下表格验证准确率显著高于图表,差距持续存在。
  • 不同模型家族的路由失效机制各异,揭示架构深层问题。

多模态大模型在辅助科学同行评审中日益重要,核心任务是判断论文中的主张是否得到证据支持。已有研究表明,当证据为表格时模型表现远优于相同数据的图表。这引发疑问:是模型无法从图表中提取信息,还是提取后未能用于决策?我们通过在三个开源视觉语言模型上进行逐层线性探针与注意力分析,研究了表格与图表证据下的表现差异。结果一致显示,图表信息虽被编码进模型中间表示,却未能抵达预测阶段——这一“信息断路”现象在表格中不存在,且在所有测试条件下均成立。注意力分析进一步揭示,该断路在不同模型族中以两种不同的架构方式呈现。这些发现将‘表格-图表差距’重新定义为预测时刻的信息路由失败,而非编码能力不足。

原文摘要 · Abstract (English)

Multimodal LLMs are increasingly used to assist scientific peer review, where a core requirement is verifying whether claims in a paper are supported by its evidence. Prior work has shown that models perform substantially better at this task when the evidence is a table than when it is a chart of the same underlying data. This raises the question of whether models fail to extract information from charts, or do they extract it but fail to use it when forming their prediction? We study this question through layer-wise linear probing and attention analysis on three open-weight VLMs over table and chart evidence, representing the same underlying data. We find consistent evidence for the latter. Chart information is encoded in the models' intermediate representations but does not reach the prediction position, a gap that is absent for tables and holds across all conditions tested. Attention analysis further reveals that this disconnect takes two architecturally distinct forms across model families. These findings reframe the table-chart gap as a failure of how encoded visual information is routed at prediction time, rather than a failure of encoding itself.

科学验证多模态模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。