arXiv:2505.17503cs.CL2025-05被引 1

构建首个评估结构化文档复杂推理的综合基准,检验大模型真实场景下的表现。

CReSt: A Comprehensive Benchmark for Retrieval-Augmented Generation with Complex Reasoning over Structured Documents

  • 设计覆盖复杂推理、引用准确性和布局理解的多维度评估框架
  • 包含2245个中英双语标注样本,模拟真实RAG应用需求
  • 适合研究检索增强生成、可信AI和大模型评测的团队使用

近年来大型语言模型在多个领域取得显著进展,但其在实际检索增强生成(RAG)场景中的能力评估仍具挑战性。在真实应用中,大模型需具备复杂推理、合理拒绝回答、精准引用和有效理解文档布局等能力,这些能力对任务处理、不确定性感知、可靠性维持和结构理解至关重要。尽管已有研究分别关注部分特性,但缺乏统一框架对这些能力进行协同评估。为此,我们提出CReSt(面向结构化文档复杂推理的检索增强生成综合基准),包含2,245个中英文人工标注样本,旨在全面捕捉需要复杂推理的实用RAG场景。该基准还引入定制化评估方法,系统衡量模型在关键维度的表现。评估结果显示,即使先进大模型在各维度上也难以保持一致表现,暴露出亟待改进的关键短板。我们已公开CReSt数据集与代码,以支持更鲁棒RAG系统的研发。数据与代码地址:https://github.com/UpstageAI/CReSt。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have made substantial progress in recent years, yet evaluating their capabilities in practical Retrieval-Augmented Generation (RAG) scenarios remains challenging. In practical applications, LLMs must demonstrate complex reasoning, refuse to answer appropriately, provide precise citations, and effectively understand document layout. These capabilities are crucial for advanced task handling, uncertainty awareness, maintaining reliability, and structural understanding. While some of the prior works address these aspects individually, there is a need for a unified framework that evaluates them collectively in practical RAG scenarios. To address this, we present CReSt (A Comprehensive Benchmark for Retrieval-Augmented Generation with Complex Reasoning over Structured Documents), a benchmark designed to assess these key dimensions holistically. CReSt comprises 2,245 human-annotated examples in English and Korean, designed to capture practical RAG scenarios that require complex reasoning over structured documents. It also introduces a tailored evaluation methodology to comprehensively assess model performance in these critical areas. Our evaluation shows that even advanced LLMs struggle to perform consistently across these dimensions, underscoring key areas for improvement. We release CReSt to support further research and the development of more robust RAG systems. The dataset and code are available at: https://github.com/UpstageAI/CReSt.

检索增强复杂推理评估基准结构化文档

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。