将网络指南转化为可自我进化的能力,让智能体更高效完成复杂任务。
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?

- 构建闭环框架,将杂乱指南编译为可编辑技能
- 在六种模型上提升12.8%至25.3%的性能
- 适合研究智能体自主学习与长周期任务的学者
网络上丰富的过程性知识蕴含巨大潜力,可帮助智能体完成长周期任务。但这些知识多模态、异构、嘈杂,且默认由人类执行,难以直接作为智能体的能力使用。为此,我们提出“指南到能力学习”问题:将真实世界指南转化为可执行技能,并基于智能体可观测轨迹持续优化。为评估现有智能体在此任务上的表现,我们引入首个相关基准MMG2Skill-Bench。进一步提出MMG2Skill框架,将指南编译为可编辑技能,用固定视觉语言模型(VLM)智能体执行,并通过轨迹级根因反馈修订技能,无需依赖基准分数。在图形界面控制、开放游戏和策略卡牌游戏中,六种VLM骨干模型均表现出色,宏观平均性能提升12.8%至25.3个百分点。消融实验表明,直接用原始指南提示会降低性能,而结构化技能构建与轨迹驱动修订缺一不可。在可成功推断的任务中,分析器式早停机制可防止后期性能下降,且在成功信号校准良好时减少25%-53%的尝试次数。
原文摘要 · Abstract (English)
Abundant procedural knowledge on the Web holds great potential for helping agents solve long-horizon tasks. However, such knowledge is often multimodal, heterogeneous, noisy, and implicitly assumes human executors, making it difficult to use directly as the skills required by agents. To bridge the gap between human-oriented guides and agent-executable skills, we formalize this problem as guide-to-skill learning: converting in-the-wild guides into executable skills and continuously improving them from trajectories observable to the agent. To evaluate the capability of existing agents on this task, we introduce MMG2Skill-Bench, the first benchmark designed for this problem. We further propose MMG2Skill, a closed-loop framework that compiles guides into editable skills, conditions a fixed vision-language model (VLM) agent on these skills during execution, and revises the skills from trajectory-level root-cause feedback without using benchmark scores. Across GUI control, open-ended gameplay, and strategic card play with six VLM backbones, MMG2Skill consistently outperforms vanilla baseline agents in every model-domain setting, achieving macro-average gains of +12.8 to +25.3 percentage points across backbones. Ablation studies show that directly prompting agents with raw guides can degrade performance, while both structured skill construction and trajectory-driven revision are necessary for the observed improvements. On success-inferable tasks, analyzer-based early stopping further prevents late-stage performance regressions and saves 25%-53% of attempts when the success signal is properly calibrated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。