arXiv:2602.03468cs.AIcs.LG2026-02

让研究型AI先搞清用户意图再行动,提升长时任务效率

IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning

  • 用强化学习训练智能体主动澄清用户潜在需求
  • 意图识别准确率提升,任务完成效果优于现有方案
  • 适合需要精准理解复杂需求的研究类应用

深度研究(DR)智能体通过自主从大规模网络语料中检索并整合证据,生成长篇报告,突破了大语言模型的参数化知识限制,实现了长期任务的自主执行。然而,与实时对话助手不同,深度研究计算成本高、耗时长,在用户意图模糊时,高自主性常导致长时间无效运行。为此,我们提出IntentRL框架,训练主动澄清隐含用户意图的智能体。为解决开放式研究数据稀缺问题,我们构建了一条可扩展的流水线,通过浅层到深层的意图细化图,将少量种子样本扩展为高质量对话片段。进一步采用两阶段强化学习策略:第一阶段在离线对话上训练通用交互行为;第二阶段利用训练好的智能体和用户模拟器进行在线试运行,增强对多样反馈的适应能力。大量实验表明,IntentRL显著提升了意图命中率和下游任务表现,优于闭源研究智能体的内置澄清模块及主动式大模型基线。

原文摘要 · Abstract (English)

Deep Research (DR) agents extend Large Language Models (LLMs) beyond parametric knowledge by autonomously retrieving and synthesizing evidence from large web corpora into long-form reports, enabling a long-horizon agentic paradigm. However, unlike real-time conversational assistants, DR is computationally expensive and time-consuming, creating an autonomy-interaction dilemma: high autonomy on ambiguous user queries often leads to prolonged execution with unsatisfactory outcomes. To address this, we propose IntentRL, a framework that trains proactive agents to clarify latent user intents before starting long-horizon research. To overcome the scarcity of open-ended research data, we introduce a scalable pipeline that expands a few seed samples into high-quality dialogue turns via a shallow-to-deep intent refinement graph. We further adopt a two-stage reinforcement learning (RL) strategy: Stage I applies RL on offline dialogues to efficiently learn general user-interaction behavior, while Stage II uses the trained agent and a user simulator for online rollouts to strengthen adaptation to diverse user feedback. Extensive experiments show that IntentRL significantly improves both intent hit rate and downstream task performance, outperforming the built-in clarify modules of closed-source DR agents and proactive LLM baselines.

深度研究强化学习意图识别智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。