arXiv:2606.20363cs.AI2026-06被引 1

从用户操作轨迹中自动挖掘可读技能库,但现有方法难提升跨领域智能体性能。

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining

论文配图:Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining
图 1 · 摘自论文原文
  • 通过三阶段流程:轨迹分段、聚类生成候选技能、训练技能感知策略
  • 五类聚类纯度超0.95,但跨域性能仅提升2个百分点,关键指标反而低于基线
  • 适合研究技能结构可解释性,不推荐用于实际策略优化

显式技能库能提升计算机使用智能体的可解释性,但其能否从交互数据中有效挖掘以改善下游策略仍不明确。本文提出三阶段管道:分割GUI轨迹、聚类段落生成候选技能、基于标注训练技能感知策略。挖掘出的聚类在InteraSkill Workflows标签下五类纯度超过0.95。然而,可读性不等于可迁移:GRPO仅将任务步准确率从18.5%提升至20.5%,BrowseComp+基本不变,且在关键源域指标上劣于简单频率先验。因此,本方法定位为诊断性研究:轨迹挖掘可揭示可解释技能结构,但当前边界检测器、无序段表示与离线奖励模型不足以支撑可靠跨域策略改进。

原文摘要 · Abstract (English)

Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from interaction data in a way that improves downstream policies. We study this question through a three-stage pipeline that segments GUI trajectories, clusters segments into candidate skills, and trains a skill-aware policy from the resulting annotations. The mined clusters are readable on the source benchmark: five of eight clusters have at least 0.95 purity against InteraSkill Workflows labels. However, readability does not imply transfer. GRPO improves IW skill-step accuracy only from 18.5\% to 20.5\%, leaves BrowseComp+ essentially unchanged, and underperforms trivial frequency priors on key source-domain metrics. We therefore present the method as a diagnostic study: trajectory mining can expose inspectable skill structure, but the current boundary detector, orderless segment representation, and offline reward model are insufficient for reliable cross-domain policy improvement.

技能挖掘可解释性轨迹分析智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。