让深度研究过程可交互可控,用户能中途调整方向。
An Interactive Paradigm for Deep Research

- 引入成本收益判断机制,动态决定是否暂停等待用户输入。
- 在多项指标上超越现有模型,对齐度提升22.8%,85%以上用户更偏好。
- 支持长期会话中的角色演化,适合需要灵活调整的研究任务。
大型语言模型的进展使得深度研究系统能够通过检索、推理与生成融合,对开放性问题给出综合性报告式回答。然而,现有框架多采用固定流程,一次设定后长时间自主运行,一旦用户意图改变则难以修正。本文提出SteER框架,实现可调控的深度研究:在每个决策点,基于成本-收益分析判断是否暂停以获取用户反馈;结合多样性感知规划与对齐、新颖性、覆盖度的效用信号,并维护动态演化的会话角色模型。实验显示,SteER在对齐度上优于主流开源及专有基线高达22.80%,在广度和均衡性等质量指标上领先,且在85%以上的成对判断中更受人类读者青睐。我们还构建了角色-查询基准与数据生成管道。据我们所知,这是首个以可交互、可解释控制范式推进深度研究的工作,为长时序任务中的可控、用户对齐智能体铺平道路。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended queries by combining retrieval, reasoning, and generation. Yet most frameworks rely on rigid workflows with one-shot scoping and long autonomous runs, offering little room for course correction if user intent shifts mid-process. We present SteER, a framework for Steerable deEp Research that introduces interpretable, mid-process control into long-horizon research workflows. At each decision point, SteER uses a cost-benefit formulation to determine whether to pause for user input or to proceed autonomously. It combines diversity-aware planning with utility signals that reward alignment, novelty, and coverage, and maintains a live persona model that evolves throughout the session. SteER outperforms state-of-the-art open-source and proprietary baselines by up to 22.80\% on alignment, leads on quality metrics such as breadth and balance, and is preferred by human readers in 85\%+ of pairwise alignment judgments. We also introduce a persona-query benchmark and data-generation pipeline. To our knowledge, this is the first work to advance deep research with an interactive, interpretable control paradigm, paving the way for controllable, user-aligned agents in long-form tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。