arXiv:2603.20133cs.CL2026-03ACL被引 2

对话场景让大模型推理能力显著下降,暴露了现有评测的局限性。

Reasoning Gets Harder for LLMs Inside A Dialogue

  • 构建动态基准BOULDER,对比孤立任务与对话中的推理表现。
  • 八款大模型在对话中平均性能下降超过30%,多轮交互是主因。
  • 适合关注模型真实对话推理能力的研究者和开发者。

大型语言模型在各类推理基准上表现优异,但这些评估通常针对孤立任务,与实际任务导向对话(TOD)场景差异显著。在TOD中,模型需在生成文本的同时完成推理,并遵循角色、格式和风格等指令。这种差异引发担忧:基准表现是否真实反映模型在对话场景中的推理鲁棒性?我们通过引入新基准BOULDER,涵盖八项旅行相关任务,涉及算术、空间与时间推理,兼具常识与形式化特征。每个问题均设置孤立与对话两种变体,实现可控对比并避免数据污染。对八款LLM的实验显示,对话设置下性能出现显著且一致的下降。消融与定性分析表明,该差距主要由多轮对话特性驱动,角色设定与工具使用要求亦有附加影响。结果强调需在真实交互场景中评估大模型的推理能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in task-oriented dialogue (TOD). In this setting, LLMs must perform reasoning inherently while generating text and adhering to instructions on role, format, and style. This mismatch raises concerns about whether benchmark performance accurately reflects models' reasoning robustness in TOD setting. We investigate how framing reasoning tasks within TOD affects LLM performance by introducing BOULDER, a new dynamic benchmark covering eight travel-related tasks that require arithmetic, spatial, and temporal reasoning with both commonsense and formal aspects. Each problem is presented in both isolated and dialogue-based variants, enabling controlled comparison while mitigating data contamination. Experiments on eight LLMs reveal a substantial and consistent performance gap between isolated and dialogue settings. Through ablations and qualitative analysis, we show that this gap is largely driven by the multi-turn nature of dialogue, with additional effects from role conditioning and tool-use requirements. Our results highlight the need to evaluate LLM reasoning in realistic interactive scenarios.

大模型推理对话系统评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。