用知识引导生成合成数据,构建多文档推理评测基准
MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance
- 基于结构化知识生成多文档,通过LLM改写引入推理挑战
- 现有模型在短文档集上仍表现不佳,验证了任务难度
- 可快速适配新挑战,适合评估多文档推理能力
自然语言处理评估得益于大语言模型(LLMs)的快速发展。随着LLMs推理能力迅速提升,构建新的评估基准成为迫切需求。特别是多文档(MD)推理因其与长上下文处理能力高度相关,但目前缺乏严谨的评测基准。由于标注长文本成本高昂,传统方法难以推进。本文提出MDBench,一个基于知识引导的合成多文档推理数据集。该方法以结构化种子知识为基础,通过LLM辅助修改生成复杂推理场景,并转换为自然文本形式,实现高效可控的文档与问答对生成。我们评估主流LLMs及提示技术,发现即使在较短文档集下,模型仍面临显著挑战。此外,该生成方法支持针对性分析推理能力,并可快速适应新挑战和模型改进。
原文摘要 · Abstract (English)
Natural language processing evaluation has made significant progress, largely driven by the proliferation of powerful large language mod-els (LLMs). New evaluation benchmarks are of increasing priority as the reasoning capabilities of LLMs are expanding at a rapid pace. In particular, while multi-document (MD) reasoning is an area of extreme relevance given LLM capabilities in handling longer-context inputs, few benchmarks exist to rigorously examine model behavior in this setting. Moreover, the multi-document setting is historically challenging for benchmark creation due to the expensive cost of annotating long inputs. In this work, we introduce MDBench, a new dataset for evaluating LLMs on the task of multi-document reasoning. Notably, MDBench is created through a novel synthetic generation process, allowing us to controllably and efficiently generate challenging document sets and the corresponding question-answer (QA) examples. Our novel technique operates on condensed structured seed knowledge, modifying it through LLM-assisted edits to induce MD-specific reasoning challenges. We then convert this structured knowledge into a natural text surface form, generating a document set and corresponding QA example. We analyze the behavior of popular LLMs and prompting techniques, finding that MDBENCH poses significant challenges for all methods, even with relatively short document sets. We also see our knowledge-guided generation technique (1) allows us to readily perform targeted analysis of MD-specific reasoning capabilities and (2) can be adapted quickly to account for new challenges and future modeling improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。