arXiv:2606.20636cs.AIcs.CL2026-06被引 2

让电脑操作智能体在动态环境中安全学习并使用技能

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

论文配图:SkillHarness: Harnessing Safe Skills for Computer-Use Agents
图 1 · 摘自论文原文
  • 通过多源监督信号识别安全技能,构建动态安全约束
  • 实验显示技能不安全率降低57.1%,执行更稳定
  • 适合需要高可靠性的自动化交互场景

计算机使用智能体(CUAs)在动态交互环境中日益普及,亟需在交互中持续学习新技能。现有方法虽能从成功轨迹中学习可复用技能,但大多假设环境静态且安全,忽视了对抗性交互(如提示注入)和环境动态性(如弹窗)带来的风险。这可能导致危险技能学习与脆弱执行,损害CUAs的可靠性。为此,我们提出SkillHarness框架,实现动态环境中的安全技能挖掘。该框架将技能学习与使用建模为受安全约束的交互过程:引入技能边界,利用多源监督信号从交互轨迹中识别安全技能,并在技能全生命周期内构建自提升安全约束;同时提出选择性技能复用机制,根据上下文分解任务并激活子集技能。实验表明,SkillHarness使学习技能的不安全率降低57.1%,并在环境动态变化下显著提升执行稳定性,优于现有基线。

原文摘要 · Abstract (English)

Computer-Use Agents (CUAs) are increasingly deployed in dynamic interactive environments, creating a growing need for continual skill learning during interaction. Recent approaches address this challenge by learning reusable skills from successful trajectories. However, these skill learning methods largely assume static and safe environments, overlooking risks from adversarial interactions (e.g., prompt injections) and environmental dynamics (e.g., pop-ups). In dynamic settings, such assumptions can lead to risky skill learning and brittle execution, undermining the reliability of CUAs. This raises the question: how can CUAs learn and use skills safely in dynamic environments? To address this problem, we propose SkillHarness, a framework for safe skill harnessing in dynamic environments. SkillHarness moves beyond static skill abstractions by modeling skill learning and utilization as a safety-constrained interaction process. Specifically, we introduce the skill boundary that leverages multi-source supervision signals to identify safe skills from interaction trajectories, and construct self-improving safety constraints throughout the skill lifecycle. In addition, SkillHarness introduces selective skill reuse, where tasks are guided to decompose according to context and completed through the selective activation of skill subsets. Our experiments demonstrate that SkillHarness significantly reduces the unsafe rate of learned skills by 57.1% and consistently improves execution stability under dynamic environmental changes, outperforming existing baselines.

智能体安全学习技能复用动态环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。