测试多模态大模型在碎片文档重建中的语义推理能力
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction

- 用Markdown自动生成碎片化文档,避免数据污染
- 碎片越多(8/12/16片),重建误差越大,性能显著下降
- 揭示当前模型缺乏跨模态细粒度推理能力
多模态大语言模型(MLLMs)在视觉丰富文档理解(VRDU)任务中表现优异,但现有评估主要针对完整、结构良好的文档图像。本文提出一种新挑战:从撕碎片段中恢复内容,需结合视觉模式识别与语义推理。为此,我们构建了ShredBench基准,基于自动化生成管道直接从Markdown渲染碎片文档,支持灵活引入最新或未见文本源,防止训练数据泄露。该基准涵盖英文、中文、代码、表格四种场景,三种碎片粒度(8、12、16片)。对主流MLLMs的实证评估显示,模型在完整文档上有效,但碎片化后重建性能急剧下降,尤其随着碎片数量增加,归一化编辑距离(NED)显著上升。结果表明,当前MLLMs缺乏弥合视觉断点所需的精细跨模态推理能力,暴露了鲁棒性文档理解研究的关键缺口。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable performance in Visually Rich Document Understanding (VRDU) tasks, but their capabilities are mainly evaluated on pristine, well-structured document images. We consider content restoration from shredded fragments, a challenging VRDU setting that requires integrating visual pattern recognition with semantic reasoning under significant content discontinuities. To facilitate systematic evaluation of complex VRDU tasks, we introduce ShredBench, a benchmark supported by an automated generation pipeline that renders fragmented documents directly from Markdown. The proposed pipeline ensures evaluation validity by allowing the flexible integration of latest or unseen textual sources to prevent training data contamination. ShredBench assesses four scenarios (English, Chinese, Code, Table) with three fragmentation granularities (8, 12, 16 pieces). Empirical evaluations on state-of-the-art MLLMs reveal a significant performance gap: The method is effective on intact documents; however, once the document is shredded, restoration becomes a significant challenge, with NED dropping sharply as fragmentation increases. Our findings highlight that current MLLMs lack the fine-grained cross-modal reasoning required to bridge visual discontinuities, identifying a critical gap in robust VRDU research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。