构建科学问答基准,要求系统精准定位证据并验证答案。
LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering
- 分三步完成科学问题回答:找论文、定证据、出答案。
- 含55个开发题例,涵盖单篇与多篇论文问题。
- 适合研发可验证科学问答系统的研究人员使用。
科学文献正被广泛用于语言模型、检索增强生成系统和研究助手,但回答科研问题不仅需要流畅生成。可靠系统需识别相关论文、定位支持答案的具体证据,并生成忠实于证据的回答。我们提出LitTraceQA,一个面向科学论文的文献支撑问答基准。给定研究问题和论文元数据池,系统须输出三个关联结果:标准论文标识符、支持证据位置,以及以自由文本、选择题或结构化表格等形式返回的答案。该基准覆盖科学阅读中常见的证据类型:表格、图表、文本片段、公式或算法、引用上下文。公开开发集包含55个例子,包括26个隐藏来源的单篇论文问题和29个多篇论文问题,提供金标论文、证据标注和答案,便于本地验证。我们还分析了更大的最终标注集,涵盖4,978个唯一问题记录,涉及4,859个唯一金标论文。通过分别评估论文检索、证据定位和答案准确性,LitTraceQA为生成可验证答案的科学问答系统提供了测试平台。
原文摘要 · Abstract (English)
Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, but answering research questions from papers requires more than fluent generation. A reliable system must identify the relevant papers, locate the concrete evidence that supports the answer, and produce a response that is faithful to that evidence. We present LitTraceQA, a benchmark for literature-grounded question answering over scientific papers. Given a research question and a metadata pool of papers, a system must return three connected outputs: canonical paper identifiers, supporting evidence locations, and answers in one or more requested formats, including free-form text, multiple-choice answers, and structured tables. LitTraceQA targets evidence types common in scientific reading: tables, figures, text spans, equations or algorithms, and citation contexts. The public development split contains 55 examples, including 26 hidden-source single-paper questions and 29 multi-paper questions, and provides gold papers, evidence annotations, and answers for local validation. We also analyze a larger final annotation collection with 4,978 unique-question records over 4,859 unique gold papers. By evaluating paper retrieval, evidence grounding, and answer accuracy separately, LitTraceQA provides a testbed for scientific QA systems that produce verifiable answers rather than unsupported summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。