arXiv:2601.09382cs.AIcs.CL2026-01被引 1

让智能代理主动追踪用户长期意图,动态响应环境变化。

Long-term Task-oriented Agent: Proactive Long-term Intent Maintenance in Dynamic Environments

  • 构建可自动生成触发条件的意图监控机制
  • 在复杂任务中实现85.19%的任务完成率,优于现有模型
  • 适合长期任务规划与动态环境交互的研究者

当前大语言模型代理多采用被动响应模式,仅在短期会话中回应即时请求,难以维持用户长期意图并适应动态环境。本文提出一种主动式任务导向智能体的新交互范式,通过两项核心能力实现主动性:(i) 意图驱动的监控——基于对话历史自主生成触发条件;(ii) 事件驱动的跟进——在检测到环境更新时主动联系用户。我们设计了高质量数据合成流程,构建复杂多轮、动态环境下的对话数据。同时,为填补动态环境下任务导向交互评估标准的空白,提出新基准ChronosBench。对主流闭源与开源模型的评测揭示其在长期任务交互中的缺陷。基于合成数据微调的模型在包含意图转移的复杂任务中达到85.19%的任务完成率,验证了数据驱动策略的有效性。

原文摘要 · Abstract (English)

Current large language model agents predominantly operate under a reactive paradigm, responding only to immediate user queries within short-term sessions. This limitation hinders their ability to maintain long-term user's intents and dynamically adapt to evolving external environments. In this paper, we propose a novel interaction paradigm for proactive Task-oriented Agents capable of bridging the gap between relatively static user's needs and a dynamic environment. We formalize proactivity through two key capabilities, (i) Intent-Conditioned Monitoring: The agent autonomously formulates trigger conditions based on dialog history; (ii) Event-Triggered Follow-up: The agent actively engages the user upon detecting useful environmental updates. We introduce a high-quality data synthesis pipeline to construct complex, multi-turn dialog data in a dynamic environment. Furthermore, we attempt to address the lack of evaluation criteria of task-oriented interaction in a dynamic environment by proposing a new benchmark, namely ChronosBench. We evaluated some leading close-source and open-source models at present and revealed their flaws in long-term task-oriented interaction. Furthermore, our fine-tuned model trained using synthetic data for supervised learning achieves a task completion rate of 85.19% for complex tasks including shifts in user intent, outperforming other models under test. And the result validated the effectiveness of our data-driven strategy.

智能代理长期任务动态环境意图保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。