分析多页图文文档理解失败原因,提出可验证的改进方向。
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

- 分离三种错误模式:表示、选择与推理,逐个测试影响。
- 缺失页面严重限制准确率,干扰项影响小,模型难跨页整合证据。
- 提示工程能改变推理行为,适合资源受限场景优化系统设计。
多页图文文档理解(MP-VRDU)需处理分散且超出模型上下文窗口的稀疏证据。现有工作对系统构建存在竞争性但未经验证的主张。本文将错误归因于表示、选择和推理三类失败模式,并通过干预单一因素而固定其余因素,在一个多页文档理解数据集上进行隔离分析。结果表明:视觉信息虽必要,但无法替代文本提取;缺失页面显著限制准确率,而干扰项影响较小;即使证据完整提供,推理模块仍无法有效整合跨页信息。提示工程可显著改变推理行为,部分提升性能但可能牺牲其他指标。基于这些发现,我们为在固定算力预算下构建此类系统提供了可操作的指导。
原文摘要 · Abstract (English)
Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has produced competing, largely untested claims about how these systems should be built. We attribute incorrect answers to three failure modes, representation, selection, and reasoning, and isolate each over a multi-page document understanding dataset by intervening on one while holding the others fixed. We find that vision is necessary but does not replace text extraction, that missing pages bound accuracy while distractors cost little, and that reasoners fail to integrate evidence across pages even when it is fully supplied. Prompting can shift reasoning behaviour substantially, improving some outcomes at the expense of others. We translate these findings into guidance for building such systems under a fixed compute budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。