arXiv:2603.01423cs.CL2026-03中稿 · the Workshop on As…被引 1

测试大模型在多轮对话中的可靠性,发现小模型退化严重。

Quantifying Conversational Reliability of Large Language Models under Multi-Turn Interaction

  • 设计三类任务模拟真实对话挑战,对比单轮与多轮表现。
  • 多轮对话下模型可靠性显著下降,小模型降幅超50%。
  • 揭示指令漂移、意图混淆等常见失败模式,适合系统评估者参考。

大型语言模型日益应用于需用户进行长时间、跨话题对话的实际场景,但其在真实多轮交互下的可靠性仍不明确。本文通过三项代表性任务系统评估对话可靠性:(1)在话题切换中保持全局约束;(2)在交错意图中正确选择工具或代理;(3)在修改和干扰下追踪结构化实体。每项任务对比单轮与多轮设置,量化对话延长带来的可靠性下降。涵盖商业与开源模型的实验显示,可靠性普遍下降,尤其小型模型更明显。错误分析揭示了指令漂移、意图混淆、上下文覆盖等重复性失效模式,损害实际系统中的可靠行为。研究强调需对大模型进行对话可靠性压力测试,并发展更鲁棒的评估方法以实现可信部署。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in real-world applications where users engage in extended, mixed-topic conversations that depend on prior context. Yet, their reliability under realistic multi-turn interactions remains poorly understood. We conduct a systematic evaluation of conversational reliability through three representative tasks that reflect practical interaction challenges: (1) maintaining global constraints across topic shifts, (2) selecting the correct tool or agent amid interleaved intents, and (3) tracking structured entities under revisions and distractions. Each task pairs single-turn and multi-turn settings, allowing us to quantify reliability degradation under extended dialogue. Across both commercial and open-source models, we observe substantial declines in reliability, particularly for smaller models. Error analyses reveal recurring failure modes such as instruction drift, intent confusion, and contextual overwriting, which compromise dependable behavior in operational systems. Our findings highlight the need for stress-testing LLMs for conversational reliability and developing more robust evaluation methods for trustworthy deployment.

大模型对话可靠性多轮交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。