arXiv:2504.10258cs.CVcs.MM2025-04被引 1

提出XY-Cut++方法,提升复杂文档布局排序准确率。

XY-Cut++: Advanced Layout Ordering via Hierarchical Mask Mechanism on a Novel Benchmark

  • 通过分层掩码预处理与多粒度分割,优化跨模态匹配。
  • 在新基准DocBench-100上达98.8 BLEU,较基线最高提升24%。
  • 适合需高精度文档结构恢复的RAG与LLM预处理场景。

文档阅读顺序恢复是文档图像理解中的基础任务,在提升检索增强生成(RAG)效果及作为大语言模型(LLMs)预处理步骤方面具有关键作用。现有方法在复杂布局(如多栏报纸)、跨模态元素间高开销交互以及缺乏稳健评估基准方面存在挑战。我们提出XY-Cut++,一种先进的布局排序方法,融合预掩码处理、多粒度分割与跨模态匹配以应对这些问题。该方法显著提升了布局排序准确性,相较于传统XY-Cut技术有明显改进。具体而言,XY-Cut++在新提出的DocBench-100数据集上实现98.8 BLEU的整体性能,较现有基线最高提升24%,并在简单与复杂布局中均表现出一致的高精度。这一进展为文档结构恢复建立了可靠基础,确立了布局排序任务的新标准,推动了更高效的RAG与LLM预处理。

原文摘要 · Abstract (English)

Document Reading Order Recovery is a fundamental task in document image understanding, playing a pivotal role in enhancing Retrieval-Augmented Generation (RAG) and serving as a critical preprocessing step for large language models (LLMs). Existing methods often struggle with complex layouts(e.g., multi-column newspapers), high-overhead interactions between cross-modal elements (visual regions and textual semantics), and a lack of robust evaluation benchmarks. We introduce XY-Cut++, an advanced layout ordering method that integrates pre-mask processing, multi-granularity segmentation, and cross-modal matching to address these challenges. Our method significantly enhances layout ordering accuracy compared to traditional XY-Cut techniques. Specifically, XY-Cut++ achieves state-of-the-art performance (98.8 BLEU overall) while maintaining simplicity and efficiency. It outperforms existing baselines by up to 24\% and demonstrates consistent accuracy across simple and complex layouts on the newly introduced DocBench-100 dataset. This advancement establishes a reliable foundation for document structure recovery, setting a new standard for layout ordering tasks and facilitating more effective RAG and LLM preprocessing.

文档理解布局排序RAG多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。