arXiv:2409.02465cs.CL2024-09被引 18

用侦探小说测试大模型长文本推理能力,发现其找证据仍困难。

DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels

  • 用平均超10万词的侦探小说构建双语问答数据集
  • 评测显示主流模型在长文本中仍难准确提取证据
  • 提出分步推理评估法,适合研究长上下文推理者

为推动大语言模型(LLMs)长上下文推理研究,我们提出DetectiveQA,一个专用于长篇叙事推理的数据集。该数据集基于平均超过10万词的侦探小说,包含1200个中英文双语人工标注的问题,每个问题配有对应的参考推理步骤。我们引入一种分步推理评估指标,以更细致地衡量模型推理过程。通过验证该方法并评测GPT-4、Claude和LLaMA等主流模型,发现其在长上下文推理中仍存在持续性挑战,尤其在证据检索方面表现不佳。研究结果为长上下文推理研究提供了重要洞见,并奠定了更严谨评估的基础。

原文摘要 · Abstract (English)

Recently, significant efforts have been devoted to enhancing the long-context capabilities of Large Language Models (LLMs), particularly in long-context reasoning. To facilitate this research, we propose \textbf{DetectiveQA}, a dataset specifically designed for narrative reasoning within long contexts. We leverage detective novels, averaging over 100k tokens, to create a dataset containing 1200 human-annotated questions in both Chinese and English, each paired with corresponding reference reasoning steps. Furthermore, we introduce a step-wise reasoning metric, which enhances the evaluation of LLMs' reasoning processes. We validate our approach and evaluate the mainstream LLMs, including GPT-4, Claude, and LLaMA, revealing persistent long-context reasoning challenges and demonstrating their evidence-retrieval challenges. Our findings offer valuable insights into the study of long-context reasoning and lay the base for more rigorous evaluations.

长上下文推理评估问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。