arXiv:2604.08618cs.IRcs.AI2026-04中稿 · ACM SIGIR 2026 Ind…被引 31

让AI客服技能自动进化,持续优化云技术支持能力。

SkillForge: Forging Domain-Specific, Self-Evolving Agent Skills in Cloud Technical Support

  • 基于真实工单和知识库生成精准技能,确保初始质量。
  • 通过失败分析自动定位问题,迭代优化技能表现。
  • 在5个真实场景中验证,比人工经验更有效提升性能。

在企业级场景如云技术支援中部署大模型驱动的智能体,需要高质量、领域专用的技能。然而现有技能生成方法缺乏领域依据,导致技能与实际任务需求脱节。此外,部署后缺乏系统性机制将执行失败追溯至技能缺陷并驱动针对性改进,致使技能质量停滞不前。本文提出SkillForge,一个闭环自演进框架,涵盖技能创建-评估-优化全过程。通过领域上下文化技能生成器,将技能合成基于知识库和历史支持工单,以获得高质量初始技能;再通过三阶段流水线——失败分析器、技能诊断器、技能优化器——批量自动诊断执行失败,精确定位技能缺陷并重写修复。该循环可迭代运行,使技能随每次部署反馈持续进化。在覆盖1,883个工单、3,737项任务的5个真实云支持场景中评估表明:(1) 领域上下文化技能生成器产生的初始技能显著优于通用生成器,其一致性更高;(2) 自演进循环能从不同起点(专家撰写、领域创建、通用技能)出发,持续提升技能质量,证明自动化进化可超越人工精心设计的知识。

原文摘要 · Abstract (English)

Deploying LLM-powered agents in enterprise scenarios such as cloud technical support demands high-quality, domain-specific skills. However, existing skill creators lack domain grounding, producing skills poorly aligned with real-world task requirements. Moreover, once deployed, there is no systematic mechanism to trace execution failures back to skill deficiencies and drive targeted refinements, leaving skill quality stagnant despite accumulating operational evidence. We introduce SkillForge, a self-evolving framework that closes an end-to-end creation-evaluation-refinement loop. To produce well-aligned initial skills, a Domain-Contextualized Skill Creator grounds skill synthesis in knowledge bases and historical support tickets. To enable continuous self-optimization, a three-stage pipeline -- Failure Analyzer, Skill Diagnostician, and Skill Optimizer -- automatically diagnoses execution failures in batch, pinpoints the underlying skill deficiencies, and rewrites the skill to eliminate them. This cycle runs iteratively, allowing skills to self-improve with every round of deployment feedback. Evaluated on five real-world cloud support scenarios spanning 1,883 tickets and 3,737 tasks, experiments show that: (1) the Domain-Contextualized Skill Creator produces substantially better initial skills than the generic skill creator, as measured by consistency with expert-authored reference responses from historical tickets; and (2) the self-evolution loop progressively improves skill quality from diverse starting points (including expert-authored, domain-created, and generic skills) across successive rounds, demonstrating that automated evolution can surpass manually curated expert knowledge.

AI客服自演化云支持技能优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。