arXiv:2410.00526cs.CL2024-10被引 2

构建新基准,评估大模型在多文档操作指南中的对话问答能力

Benchmarking Large Language Models for Conversational Question Answering in Multi-instructional Documents

  • 针对多文档操作指南设计新型对话问答评测集
  • 模型需从多份说明中提取并整合步骤指引,准确率是关键指标
  • 适合研究智能助手、复杂任务理解的开发者使用

instructional documents 是完成各类任务的重要知识来源,但其在对话式问答(CQA)中的独特挑战尚未得到充分研究。现有基准主要聚焦于单篇叙述性文档的基础事实问答,难以评估模型在真实世界多文档操作指南中理解与提供逐步指导的能力。为此,我们提出 InsCoQA,一个专为评估大语言模型(LLMs)在多文档指令类文本中进行对话式问答而设计的新基准。该数据集源自大量百科式操作内容,要求模型具备从多份文档中检索、解读并准确总结操作流程的能力,体现真实任务的复杂性和多面性。此外,为全面评估前沿大模型在 InsCoQA 上的表现,我们提出 InsEval——一种基于 LLM 的辅助评估框架,用于衡量生成回答和操作指令的完整性与准确性。

原文摘要 · Abstract (English)

Instructional documents are rich sources of knowledge for completing various tasks, yet their unique challenges in conversational question answering (CQA) have not been thoroughly explored. Existing benchmarks have primarily focused on basic factual question-answering from single narrative documents, making them inadequate for assessing a model`s ability to comprehend complex real-world instructional documents and provide accurate step-by-step guidance in daily life. To bridge this gap, we present InsCoQA, a novel benchmark tailored for evaluating large language models (LLMs) in the context of CQA with instructional documents. Sourced from extensive, encyclopedia-style instructional content, InsCoQA assesses models on their ability to retrieve, interpret, and accurately summarize procedural guidance from multiple documents, reflecting the intricate and multi-faceted nature of real-world instructional tasks. Additionally, to comprehensively assess state-of-the-art LLMs on the InsCoQA benchmark, we propose InsEval, an LLM-assisted evaluator that measures the integrity and accuracy of generated responses and procedural instructions.

对话问答多文档指令理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。