测试大模型长文档推理能力,设计了100个专家级多步问题。
DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
- 通过人机协作构建真实长文档问答任务,确保难度与质量。
- o1-preview模型正确率达69.7%,显著优于通用模型的57.7%。
- 推理模型蒸馏后性能下降明显,说明知识迁移仍有挑战。
我们提出了DocPuzzle,一个严谨构建的基准,用于评估大语言模型(LLMs)在长上下文中的推理能力。该基准包含100个需基于真实世界长文档进行多步推理的专家级问答问题。为保证任务质量和复杂性,我们采用人机协同的标注-验证流程。DocPuzzle引入创新的评估框架,通过检查清单引导的过程分析,有效缓解猜测偏差,建立了评估大模型推理能力的新标准。评估结果显示:1)先进慢思考推理模型如o1-preview(69.7%)和DeepSeek-R1(66.3%)显著优于最佳通用指令模型Claude 3.5 Sonnet(57.7%);2)蒸馏后的推理模型如DeepSeek-R1-Distill-Qwen-32B(41.3%)远低于教师模型表现,表明仅靠蒸馏难以保持推理能力的泛化性。
原文摘要 · Abstract (English)
We present DocPuzzle, a rigorously constructed benchmark for evaluating long-context reasoning capabilities in large language models (LLMs). This benchmark comprises 100 expert-level QA problems requiring multi-step reasoning over long real-world documents. To ensure the task quality and complexity, we implement a human-AI collaborative annotation-validation pipeline. DocPuzzle introduces an innovative evaluation framework that mitigates guessing bias through checklist-guided process analysis, establishing new standards for assessing reasoning capacities in LLMs. Our evaluation results show that: 1)Advanced slow-thinking reasoning models like o1-preview(69.7%) and DeepSeek-R1(66.3%) significantly outperform best general instruct models like Claude 3.5 Sonnet(57.7%); 2)Distilled reasoning models like DeepSeek-R1-Distill-Qwen-32B(41.3%) falls far behind the teacher model, suggesting challenges to maintain the generalization of reasoning capabilities relying solely on distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。