arXiv:2505.06120cs.CLcs.HC2025-05被引 418

大模型在多轮对话中易迷失方向,表现大幅下降。

LLMs Get Lost In Multi-Turn Conversation

  • 通过大规模模拟对比单轮与多轮对话性能。
  • 多轮任务平均性能下降39%,主要因不可靠性上升。
  • 适合关注对话系统鲁棒性的研究人员参考。

大型语言模型(LLMs)是对话式接口,具备在用户未完全明确需求时协助其定义、探索和优化任务的能力。尽管分析显示用户指令常存在不明确情况,但现有评测仍以单轮全明确指令为主。本研究通过大规模模拟实验,比较了主流开源与闭源LLM在单轮与多轮场景下的表现。结果表明,所有测试模型在多轮对话中性能显著下降,六项生成任务平均下降39%。对20万+次模拟对话的分析将性能退化归因于两个因素:能力轻微下降和可靠性显著降低。发现模型常在早期对话中做出假设并过早生成最终答案,且过度依赖这些错误推断。简言之,一旦在对话中走错路,模型便难以恢复。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are conversational interfaces. As such, LLMs have the potential to assist their users not only when they can fully specify the task at hand, but also to help them define, explore, and refine what they need through multi-turn conversational exchange. Although analysis of LLM conversation logs has confirmed that underspecification occurs frequently in user instructions, LLM evaluation has predominantly focused on the single-turn, fully-specified instruction setting. In this work, we perform large-scale simulation experiments to compare LLM performance in single- and multi-turn settings. Our experiments confirm that all the top open- and closed-weight LLMs we test exhibit significantly lower performance in multi-turn conversations than single-turn, with an average drop of 39% across six generation tasks. Analysis of 200,000+ simulated conversations decomposes the performance degradation into two components: a minor loss in aptitude and a significant increase in unreliability. We find that LLMs often make assumptions in early turns and prematurely attempt to generate final solutions, on which they overly rely. In simpler terms, we discover that *when LLMs take a wrong turn in a conversation, they get lost and do not recover*.

对话系统大模型性能退化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。