大模型在用户意图动态变化时容易迷失,表现大幅下降。
LLMs Get Lost in Evolving User Intent

- 构建动态对话框架,让用户意图在多轮中逐步揭示和改变。
- 多种任务下模型性能显著下降,强静态表现无法迁移。
- 揭示评估盲区,对协作型智能体至关重要。
随着大模型能力增强,它们越来越多地被用作协作代理,通过多轮交互完成用户委托的任务。然而真实交互本质上是动态的:用户很少一开始就明确意图,而是在对话过程中逐步披露、修改甚至重新定向。尽管如此,大模型仍主要在单轮、完整指定的场景下进行评估或训练,这留下一个根本问题:大模型如何跟踪并响应随对话演进的用户意图?为此,我们提出一个框架,将静态单轮任务转化为动态多轮对话,使用户意图在对话中逐步揭示、修改甚至中途转向,同时保留原有评估协议,使现有基准可直接复用。在多个任务上,我们发现一致现象:在静态设置下表现优异的模型,在意图演变场景下性能大幅下降,不同模型家族均出现显著退化。研究结果揭示了一个根本性差距:当前大模型尚不能忠实追踪并响应用户不断变化的意图——这一能力虽在静态评估中不可见,却是未来协作代理的关键。
原文摘要 · Abstract (English)
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users rarely specify their intent upfront, instead disclosing, revising, and reshaping it as the conversation unfolds. Despite this, LLMs are still predominantly evaluated or trained in single-turn, fully-specified settings, leaving open a fundamental question: how well do LLMs track and act on user intent as it evolves over the course of a conversation? To study this, we introduce a framework that transforms static, single-turn tasks into dynamic multi-turn conversations in which the user's intent evolves across turns--incrementally revealed, revised, and at times redirected mid-conversation--while preserving each task's original evaluation protocol, enabling existing benchmarks to be reused as controlled testbeds without new annotation. Across multiple tasks, we surface a consistent phenomenon: strong static-setting performance does not transfer to the evolving-intent setting, with substantial drops across model families. Our findings point to a fundamental gap: today's LLMs do not yet faithfully track and act on the user's evolving intent, a capability invisible to static evaluation yet critical for future collaborative agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。