arXiv:2602.17990cs.AI2026-02

为多智能体工作流评估设计可控扰动测试,帮工程师判断更新是否安全。

WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics

  • 通过三种真实扰动(缺步骤、压缩步骤、描述变更)分级测试工作流性能
  • 4973个基准流程生成44757个扰动版本,覆盖10%~50%严重度
  • 揭示不同评估指标的敏感性差异,支持按严重度解读评分变化

从自然语言请求生成结构化工作流的多智能体大模型系统已广泛应用于云自动化、DevOps及企业流程编排。然而,日常更新(如重新运行相同输入、更换底层大模型、修改智能体提示词或编排代码)常导致生成工作流与已验证基准产生显著差异,工程师缺乏可靠依据判断变更是否可发布。自动工作流评估本是解决此问题的关键工具,但现有指标得分校准不足,数值变化难以反映实际退化程度。本文提出 WorkflowPerturb,一个通过施加真实、分级扰动于黄金流程的受控基准。该数据集包含4,973个黄金流程及44,757个扰动变体,涵盖缺失步骤、压缩步骤、描述变更三类扰动,每类在10%、30%、50%三个严重度层级上实施。我们对多种指标家族进行基准测试,并通过预期得分轨迹与残差分析其灵敏度与校准性。结果揭示了不同指标族间的系统性差异,支持在变更管理场景中进行严重度感知的评估分数解读。数据集将在论文接受后公开。

原文摘要 · Abstract (English)

Multi-agent LLM systems that generate structured workflows from natural-language requests are now deployed in production across cloud automation, DevOps, and enterprise process orchestration. Operating such systems exposes a recurring change-management problem. Routine updates, such as re-running the same input, swapping the underlying LLM, or refactoring an agent's prompt or orchestration code, frequently produce workflows that differ substantially from previously validated references. Engineers are then left without a principled way to decide whether a change is safe to ship. Automatic workflow evaluation is the natural tool for answering this question. In practice, however, metric scores are poorly calibrated, and a numeric change rarely communicates the severity of the underlying degradation. We introduce WorkflowPerturb, a controlled benchmark for studying workflow evaluation metrics by applying realistic, graded perturbations to golden workflows. WorkflowPerturb contains 4,973 golden workflows and 44,757 perturbed variants across three perturbation types (Missing Steps, Compressed Steps, and Description Changes), each applied at severity levels of 10%, 30%, and 50%. We benchmark multiple metric families and analyze their sensitivity and calibration using expected score trajectories and residuals. Our results characterize systematic differences across metric families and support severity-aware interpretation of workflow evaluation scores in change-management settings. Our dataset will be released upon acceptance.

多智能体工作流评估评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。