构建多轮对话评估基准,测试大模型的逻辑推理与信息交互能力。
Multi-Turn Puzzles: Evaluating Interactive Reasoning and Strategic Dialogue in LLMs
- 设计多轮任务,聚焦推理与主动求知能力
- 发现主流模型在指令遵循和规划上存在明显短板
- 适合研究对话智能与复杂任务处理的学者
大型语言模型(LLMs)在处理明确完整的问题时表现优异,但在现实场景中常见的模糊环境或交互式任务中常表现不佳。这凸显了发展具备逻辑一致性的多轮对话、主动获取信息并基于不完整数据进行推理能力的重要性。为此,我们提出一个全新的基准,包含一系列多轮任务,每项任务专门测试特定的推理、交互对话和信息获取能力。这些任务采用确定性评分机制,无需人工参与。对前沿模型的评估显示显著提升空间。分析表明,主要错误源于指令遵循不良、推理失败和规划不足。该基准为理解当前大模型在复杂交互场景中的优劣势提供了关键洞察,并为未来提升相关能力的研究提供了可靠平台。
原文摘要 · Abstract (English)
Large language models (LLMs) excel at solving problems with clear and complete statements, but often struggle with nuanced environments or interactive tasks which are common in most real-world scenarios. This highlights the critical need for developing LLMs that can effectively engage in logically consistent multi-turn dialogue, seek information and reason with incomplete data. To this end, we introduce a novel benchmark comprising a suite of multi-turn tasks each designed to test specific reasoning, interactive dialogue, and information-seeking abilities. These tasks have deterministic scoring mechanisms, thus eliminating the need for human intervention. Evaluating frontier models on our benchmark reveals significant headroom. Our analysis shows that most errors emerge from poor instruction following, reasoning failures, and poor planning. This benchmark provides valuable insights into the strengths and weaknesses of current LLMs in handling complex, interactive scenarios and offers a robust platform for future research aimed at improving these critical capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。