arXiv:2604.12047cs.CLcs.IR2026-04被引 1

评测不同PDF解析与分块策略对金融问答RAG系统的影响

Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG

  • 对比多种PDF解析器和分块方式,研究其对信息保留的影响
  • 在金融领域两个基准上验证,发现分块重叠提升答案准确率
  • 为构建稳健的PDF理解RAG系统提供可落地的设计建议

PDF文件主要面向人类阅读而非自动化处理,其内容异构性(如文本、表格、图像)给解析与信息提取带来挑战。为应对这些问题,从业者和研究者正开发新方法,尤其是有前景的检索增强生成(RAG)系统。然而,目前尚无全面研究探讨不同组件与设计选择如何影响RAG系统在理解PDF方面的性能。本文通过聚焦问答这一具体语言理解任务,并利用两个金融领域的基准(包括我们新构建的公开数据集TableQuest),系统评估了多种PDF解析器与分块策略(含不同重叠度),分析其在保持文档结构和确保答案正确性方面的协同效应。整体结果为构建稳健的PDF理解RAG流水线提供了实用指导。

原文摘要 · Abstract (English)

PDF files are primarily intended for human reading rather than automated processing. In addition, the heterogeneous content of PDFs, such as text, tables, and images, poses significant challenges for parsing and information extraction. To address these difficulties, both practitioners and researchers are increasingly developing new methods, including the promising Retrieval-Augmented Generation (RAG) systems to automated PDF processing. However, there is no comprehensive study investigating how different components and design choices affect the performance of a RAG system for understanding PDFs. In this paper, we propose such a study (1) by focusing on Question Answering, a specific language understanding task, and (2) by leveraging two benchmarks from the financial domain, including TableQuest, our newly generated, publicly available benchmark. We systematically examine multiple PDF parsers and chunking strategies (with varied overlap), along with their potential synergies in preserving document structure and ensuring answer correctness. Overall, our results offer practical guidelines for building robust RAG pipelines for PDF understanding.

PDF解析RAG金融问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。