arXiv:2607.25560cs.AI2026-07被引 2

通过行为轨迹逆向推断隐藏的智能体技能,揭示了隐私泄露风险。

Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories

论文配图:Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
图 1 · 摘自论文原文
  • 利用任务轨迹中的重复模式识别技能特征
  • 平均提升成功率6.88个百分点,优于基线方法
  • 适合关注AI安全与隐私保护的研究者

智能体技能以轻量可复用的形式存在,能提升下游任务表现。其便携性使其适用于市场交易和私有部署,促使提供方将高价值技能保密。然而,技能的行为影响仍可通过执行轨迹暴露,形成行为侧信道。本文定义此现象为‘技能泄漏’:仅凭良性查询生成的轨迹,无需参考答案或成功标签,即可重建专有技能。提出SigLeak框架,通过构造多样、决策丰富的诊断任务,对比启用与禁用技能的轨迹,迭代重构出技能模式。在五个场景、三种模型族和三种智能体框架中,SigLeak几乎在所有设置下均优于或匹配三个基线。平均成功率较禁用技能基准提升6.88个百分点,并在衡量粗粒度与细粒度语义相似性的SkillSim指标上达到最高。结果表明,良性执行轨迹可暴露专有程序知识。代码已公开于https://anonymous.4open.science/r/SigLeak-D1DB。

原文摘要 · Abstract (English)

Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetization and private deployment behind cloud-hosted agent interfaces, giving providers incentives to keep high-value skills proprietary. Yet hiding the artifacts does not conceal their behavioral effects, which remain observable in execution trajectories and form a behavioral side channel. We define this exposure as Skill Leakage: reconstructing proprietary skills from trajectories elicited by benign queries, without reference answers or success labels. We introduce SigLeak, a black-box framework that exploits recurring skill signatures in agent behavior. It constructs diverse, decision-rich diagnostic tasks, contrasts matched skill-enabled and skill-disabled trajectories, and iteratively refines a reconstructed skill from the isolated patterns. Across five scenarios, three model families, and three agent frameworks, SigLeak outperforms or matches three baselines in nearly every setting. It raises the success rate by 6.88 percentage points over the skill-disabled reference on average and achieves the highest overall SkillSim, our metric for coarse- and fine-grained semantic similarity. These results show that benign execution trajectories can expose proprietary procedural knowledge. The code is available at https://anonymous.4open.science/r/SigLeak-D1DB.

AI安全技能推理行为分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。