arXiv:2604.23815cs.CL2026-04被引 1

首个针对科研智能体中间动作的用户反馈数据集,揭示如何设计更懂用户意图的行动策略。

DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

论文配图:DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
图 1 · 摘自论文原文
  • 构建首个包含用户对中间动作偏好与执行效果判断的数据集
  • 发现模型需依赖完整历史行为才能准确预测用户选择,而非仅靠自我陈述
  • 提出基于历史交互生成新动作的在线干预,用户采纳率最高

科学深度研究(DR)智能体通过整合论文生成多章节报告以回答用户问题。现有评估仅关注最终报告质量,难以分析哪些中间操作能提升报告效果。我们构建了首个包含用户对中间动作反馈的数据集DRACULA:19位计算机领域专家在五周内向一个提出动作建议(如“添加数据集章节”)的系统提问,选择偏好动作,并判断报告是否成功执行其选择,共收集8,103条动作偏好和5,230条执行判断。确认智能体可执行这些动作后,我们通过模拟研究大模型预测用户偏好能力:发现(1)初始模型预测困难,但结合用户完整选择历史时表现最佳;(2)同一查询下用户选择差异源于未明示目标,阻碍模拟效果;(3)基于历史交互生成新动作的在线干预,在后续实验中被用户最频繁采纳。结果表明,当前研究聚焦于执行,而真正挑战在于决定执行哪些动作。我们开源了整个研究设计、用户反馈及模拟任务,推动长周期智能体的动作反馈研究。

原文摘要 · Abstract (English)

Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their utility, but existing protocols only score the final report, making it hard to study and learn which intermediate actions DR agents should take to improve reports. We collect DRACULA, the first dataset with user feedback on intermediate actions for DR. Over five weeks, nineteen expert CS researchers ask queries to a DR system that proposes actions (e.g., "Add a section on datasets"). Our users select actions they prefer, then judge whether an output report applied their selections successfully, yielding 8,103 action preferences and 5,230 execution judgments. After confirming a DR agent can execute DRACULA's actions, we study the predictability of user-preferred actions via simulation-how well LLMs predict the actions users select-a step toward learning to generate useful actions. We discover: (1) LLM judges initially struggle to predict action selections, but improve most when using a user's full selection history, rather than self-reported or extrapolated user context signals; (2) Users' selections for the same query differ based on unstated goals, bottlenecking simulation and motivating affordances that let users steer reports; and (3) Our simulation results inform an online intervention that generates new actions based on the user's past interactions, which users pick most often in follow-up studies. Overall, while work extensively studies execution, DRACULA reveals a key challenge is deciding which actions to execute in the first place. We open-source DRACULA's study design, user feedback, and simulation tasks to spur future work on action feedback for long-horizon agents.

智能体用户反馈科研助手动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。