arXiv:2608.22610cs.AI2026-08

提升智能体技能可靠性,避免无效或反效果的技能组合。

Coalition-Aware Skill Reliability for Self-Evolving Agents

论文配图:Coalition-Aware Skill Reliability for Self-Evolving Agents
图 1 · 摘自论文原文
  • 通过采样肖尔利边际值筛选更可靠的技能,防止污染性技能进入知识库。
  • 在跨域迁移中自动屏蔽降低检索质量的技能,提升泛化能力。
  • 适用于需要长期自主学习与稳定推理的复杂任务场景。

智能体技能是从交互轨迹中提炼出的结构化知识,可动态复用以实现大语言模型驱动的自演化智能体从经验中学习。现有研究多关注技能的获取、演化与检索等操作层面,却忽略了更根本的可靠性问题:技能库中积累的技能是否真正带来机制性正向贡献?本文通过系统性的技能库审计,在不同组合与部署领域下评估技能变化对智能体行为的影响。发现两类可靠性失效:联盟污染(银行级收益掩盖了联盟级负贡献)和跨域效用反转(源域有益技能在迁移后效果逆转)。为此提出两种干预策略:在技能积累阶段采用联盟感知的技能选择(CASS),基于采样的肖尔利边际值筛选可靠候选;在迁移后采用无标签技能掩码(u-SMCO),通过排除能提升未标注目标域检索质量的技能来优化性能。在LoCoMo、LongMemEval、HotpotQA和ALFWorld上的实验表明,CASS与u-SMCO在多个强基线之上持续提升任务表现与跨域泛化能力。此外,该方法还能降低强化学习中对噪声奖励信号的敏感性,并揭示孤立式技能评估的局限性。

原文摘要 · Abstract (English)

Agent skills, structured artifacts distilled from interaction trajectories and dynamically reused from skill banks, have become a central mechanism for enabling large language model (LLM)-based self-evolving agents to learn from past experience. Yet existing work has largely focused on the operational aspects of skills, such as acquisition, evolution, and retrieval, while leaving a more fundamental reliability question unresolved: Do accumulated skills in an agent's skill bank actually make positive mechanistic contributions? We investigate this question through systematic skill-bank audits across alternative bank compositions and deployment domains, measuring the resulting changes in agent behavior. These audits reveal two recurring reliability failures: coalition pollution, where bank-level gains conceal negative coalition-level skill contributions, and cross-domain utility reversal, where source-beneficial skills reverse their effects after transfer. These findings motivate two reliability interventions: coalition-aware skill selection during skill accumulation and label-free skill masking after transfer. Coalition-Aware Skill Selection (CASS) selects more reliable candidate skills for the current bank using sampled Shapley marginals. Unsupervised Skill-Masked Coalition Optimizer (u-SMCO) masks transferred skills whose exclusion improves retrieval quality on unlabeled target-domain data. Agentic experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld show that CASS and u-SMCO consistently improve task performance and cross-domain generalization over strong skill-based self-evolving agent baselines. Beyond accuracy, coalition-conditioned reliability modeling reduces sensitivity to noisy outcome-reward fluctuations during reinforcement learning and exposes the limits of isolation-based skill evaluation.

智能体技能可靠性自演化跨域迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。