arXiv:2505.17123cs.CL2025-05被引 14

构建首个面向多轮交互推理的自动化评估基准

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

  • 设计40个任务、3600个实例的多轮推理数据集
  • 顶尖模型在多轮交互中表现仍显著不足
  • 适合研究交互式AI与复杂推理的学者使用

大语言模型在复杂推理任务上取得进展,但现有评估多集中于单轮场景,缺乏对交互式任务的系统考察。我们提出MTR-Bench,用于大模型多轮推理评估。该基准包含4类、40个任务、3600个实例,覆盖多样化推理能力与细粒度难度层级,要求模型与环境进行多轮交互。同时,其全流程自动化框架实现无需人工干预的可扩展评估。大量实验表明,即使最先进的推理模型在多轮交互任务中仍表现不佳。对结果的深入分析为未来交互式AI系统研究提供了重要启示。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly focus on single-turn reasoning scenarios, leaving interactive tasks largely unexplored. We attribute it to the absence of comprehensive datasets and scalable automatic evaluation protocols. To fill these gaps, we present MTR-Bench for LLMs' Multi-Turn Reasoning evaluation. Comprising 4 classes, 40 tasks, and 3600 instances, MTR-Bench covers diverse reasoning capabilities, fine-grained difficulty granularity, and necessitates multi-turn interactions with the environments. Moreover, MTR-Bench features fully-automated framework spanning both dataset constructions and model evaluations, which enables scalable assessment without human interventions. Extensive experiments reveal that even the cutting-edge reasoning models fall short of multi-turn, interactive reasoning tasks. And the further analysis upon these results brings valuable insights for future research in interactive AI systems.

多轮推理评估基准交互式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。