构建首个覆盖十维复杂度的PDF问答数据集,支持真实与合成样本评估。
pdfQA: Diverse, Challenging, and Realistic Question Answering over PDFs
- 构建2000个真实与2000个合成的PDF问答对,涵盖10类复杂度维度
- 通过质量与难度过滤后,得到具有挑战性的高质量问答对
- 适合评估端到端文档问答系统在信息抽取与解析中的综合能力
PDF是互联网上使用第二广泛的文档类型(仅次于HTML)。然而,现有问答数据集通常基于文本源或仅针对特定领域。本文提出pdfQA,包含2000个由人类标注的真实问答对(real-pdfQA)和2000个合成问答对(syn-pdfQA),在十种复杂度维度(如文件类型、来源模态、位置、答案类型等)上进行区分。我们对两个数据集应用质量与难度筛选,获得有效且具挑战性的问答对。使用开源大模型回答问题,揭示了现有挑战与复杂度维度之间的关联。pdfQA为端到端问答流程评估提供了基础,可用于测试多样技能和局部优化(如信息检索或解析)。
原文摘要 · Abstract (English)
PDFs are the second-most used document type on the internet (after HTML). Yet, existing QA datasets commonly start from text sources or only address specific domains. In this paper, we present pdfQA, a multi-domain 2K human-annotated (real-pdfQA) and 2K synthetic dataset (syn-pdfQA) differentiating QA pairs in ten complexity dimensions (e.g., file type, source modality, source position, answer type). We apply and evaluate quality and difficulty filters on both datasets, obtaining valid and challenging QA pairs. We answer the questions with open-source LLMs, revealing existing challenges that correlate with our complexity dimensions. pdfQA presents a basis for end-to-end QA pipeline evaluation, testing diverse skill sets and local optimizations (e.g., in information retrieval or parsing).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。