arXiv:2606.18191cs.AIcs.MA2026-06

构建首个个性化工作流预测基准,评估智能体从多源信息中推理任务步骤的能力。

DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction

论文配图:DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction
图 1 · 摘自论文原文
  • 设计跨领域任务,要求智能体从分散来源中提取证据并推断操作序列。
  • 包含100个任务、3900+来源,7项诊断指标验证流程准确性与个性化程度。
  • 提出参考智能体DRFA,但性能仍有明显提升空间,凸显该任务的挑战性。

深度研究系统广泛应用于复杂信息检索任务,但现有研究多聚焦于生成报告与摘要。而企业实际需求更关注智能体识别具体操作流程——即一系列动作步骤。例如,面对“在固定预算下如何申请新增编制”问题,智能体需给出准确的操作序列而非仅总结政策。为此,本文提出DRFLOW,一个用于评估智能体从异构来源中预测个性化工作流的基准数据集。每个任务要求智能体从分散来源中识别相关证据,并据此预测用户任务的正确动作步骤序列。DRFLOW涵盖五个领域共100个任务,包含超过3900份来源和1246条参考流程步骤。我们定义了七项诊断指标,覆盖事实依据、步骤恢复、结构排序、条件解析及个性化等维度。进一步提出工作流导向的参考智能体DRFA(DRFLOW-Agent),实验显示其虽优于强基线模型(平均F1提升达10.02%),但在多项指标上仍存在显著提升空间,表明完整且准确的个性化工作流预测仍是深度研究的重要挑战。

原文摘要 · Abstract (English)

Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating reports and summaries. In contrast, many enterprise tasks instead require an agent to identify concrete workflows which is a sequence of action-steps. For example, rather than summarizing budgeting policies, an agent should be able to determine the steps needed to answer a question such as: "How do I request new headcount given a fixed budget?". Therefore, we introduce DRFLOW, a benchmark for evaluating personalized workflows predicted by agents from heterogeneous sources. Each task requires the agent to identify relevant evidence from scattered sources, then use that evidence to predict the correct action-step sequence for the user's task. DRFLOW contains 100 tasks across five domains, with 1,246 reference workflow steps grounded in more than 3,900 sources. We define seven diagnostic metrics covering factual grounding, step recovery, structural ordering, condition resolution, and personalization. We further present DRFLOW-Agent (DRFA), a workflow-oriented reference agent to predict personalized workflow. We show that although DRFA improves over strong baseline agents (upto 10.02% average F1 score), there is substantial room for improvement remains across these workflow metrics, indicating that predicting complete and correct personalized workflows remains a challenging frontier for deep research.

工作流预测智能体评估深度研究多源推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。