arXiv:2604.01608cs.AI2026-04被引 10

将多智能体工作流提炼为单智能体技能,关键在判断何时保留流程指导。

From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?

论文配图:From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
图 1 · 摘自论文原文
  • 区分能力资源与流程引导,识别可提炼的关键组件。
  • 引入行为-结果自由度指标F,解释性能提升与下降的反转现象。
  • 提出AdaSkill模型,按F值动态继承流程指导,降低部署开销。

针对结构化数据科学任务的多智能体系统通过跨阶段、工具、共享状态、验证与修复的工作流外化分析控制。将此类工作流提炼为单智能体技能可减少编排开销,但尚不清楚哪些组件应跨越控制边界。本文区分能力资源(扩展智能体能力)与流程引导(约束探索路径)。在相同的因果估计实例上,向能力匹配的技能添加任务限定的流程引导,在方法选择准确率下使标准化效用提升19.6点,但在数值误差下反而下降10.3点。为解释这一反转,提出行为-结果自由度(F),作为预合成诊断指标,量化行为与结果秩的符号不匹配,并通过有符号锚定秩迁移形式化其条件作用。基于此机制,提出AdaSkill:保留已验证的能力资源,移除运行时编排,根据校准后的F值条件性继承流程引导。在16次能力匹配干预中,原生规模的全减剔除效应随连续F值下降(r = -0.80, p < 0.001);15项治疗原子扫描将性能反转定位至流程引导。在11个数据集覆盖四个结构化数据科学任务类别中,AdaSkill兼具优异任务表现与显著更低的部署开销。

原文摘要 · Abstract (English)

Multi-agent systems (MAS) for structured data-science tasks externalize analytical control through workflows spanning stages, tools, shared state, verification, and repair. Distilling such workflows into a single-agent skill can reduce orchestration overhead, but it remains unclear which workflow components should cross the control boundary. We distinguish capability resources, which expand what an agent can do, from pipeline guidance, which constrains which solutions it explores. On the same causal-estimation instances, adding task-qualified source pipeline guidance to a capability-matched skill changes normalized utility by +19.6 points under method-selection accuracy but -10.3 points under numerical error. To explain this reversal, we introduce Behavior-Outcome Freedom (F), a pre-synthesis diagnostic of signed behavior-outcome rank mismatch, and formalize its candidate-conditional role through Signed Anchor-Rank Transfer. Motivated by this mechanism, we propose AdaSkill, which preserves validated capability resources, removes runtime orchestration, and conditionally inherits pipeline guidance using a calibrated rule over F. Across 16 capability-matched interventions, the native-scale Full-minus-Discard effect decreases across the continuous F scale (r = -0.80, p < 0.001), while a 15-treatment atomic sweep localizes the reversal to pipeline guidance. Across 11 datasets spanning four structured data-science task families, AdaSkill combines strong task performance with substantially lower deployment overhead.

多智能体技能提炼自动化推理部署优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。