发现视觉文档理解中模型内部表示与输出回答存在差距,中间层更易捕捉任务信息。
Responses Fall Short of Understanding: Revealing the Gap between Internal Representations and Responses in Visual Document Understanding
- 用线性探测分析LVLM各层对文档任务信息的编码能力
- 中间层比最终层更线性地编码任务所需信息,且与回答差距更小
- 针对中间层微调可同时提升表示质量和输出准确率
视觉文档理解(VDU)是大型视觉语言模型(LVLMs)面临的挑战性任务,需融合视觉感知、文本识别与结构化布局推理。尽管近期LVLM在VDU基准上取得进展,但其性能通常基于生成的回答评估,未必反映模型是否真正内化了所需信息。本文通过线性探测研究不同层中解决VDU任务的信息表示。结果表明:(1)内部表示与生成回答间存在明显差距;(2)任务相关信息在中间层比在最终层更线性编码。基于此,我们探索针对中间层的微调策略。实验显示,中间层微调能同时提升线性探测准确率与回答准确率,缩小二者差距。
原文摘要 · Abstract (English)
Visual document understanding (VDU) is a challenging task for large vision language models (LVLMs), requiring the integration of visual perception, text recognition, and reasoning over structured layouts. Although recent LVLMs have shown progress on VDU benchmarks, their performance is typically evaluated based on generated responses, which may not necessarily reflect whether the model has actually captured the required information internally. In this paper, we investigate how information required to solve VDU tasks is represented across different layers of LLMs within LVLMs using linear probing. Our study reveals that (1) there is a clear gap between internal representations and generated responses, and (2) information required to solve the task is often encoded more linearly from intermediate layers than from the final layer. Motivated by these findings, we explore fine-tuning strategies that target intermediate layers. Experiments show that fine-tuning intermediate layers improves both linear probing accuracy and response accuracy while narrowing the gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。