评估智能体技能检索中同族风险暴露问题,发现高召回率下仍易选错执行细节。
Right Family, Wrong Skill: Benchmarking Risk Exposure in Agent Skill Retrieval

- 构建同能力风险检索基准,测试技能家族内相似但执行条件不同的风险项
- 公开模型在召回率达0.848时,仍有34.6%~37.2%概率选错风险技能
- 提出用HSR@K衡量同族风险暴露,适合关注安全性的技能系统设计者
智能体技能库正成为可复用的软件资产:检索到的技能会带来指令、脚本、资源绑定和执行假设。这使得检索失败更具体而非泛化无关。系统可能找到正确能力家族,却暴露了同一能力下的错误代表。我们研究此类失败为同能力风险暴露检索。每个基准单元配对一个有用技能与一个查询相关的风险兄弟技能,二者共享能力家族但执行控制契约不同(如资源需求、前提条件、流程或产出物)。我们引入SameCapRisk-Bench,一个可审计的基准,包含1,190个技能-风险单元和1,686个评估查询案例:694个在公开库压力下的标记兄弟单元,以及496个硬角色互换单元(相同两技能在成对查询中交换有用/风险角色)。发布记录包括采纳证据、提示/泄漏检查、源哈希、家族关系和固定候选池。基准报告帮助性排名及有害兄弟比例(HSR@K),即前K位中被暴露的风险兄弟比例。在该基准上,公开的SkillRouter、SkillRet和R3-Skill在Recall@3达0.848–0.888,但频繁暴露标记风险兄弟(HSR@3 0.346–0.372)。全公开的评分聚类管道将HSR@3降至0.128–0.182,同时保持Recall@3为0.713–0.776。在基准训练的参考评分器下,公开文本聚类和受控解析器分别达到HSR@3 0.012和0.007;后者实现Recall@3 0.833。因此,技能检索应同时报告能力匹配与同族风险暴露,以HSR作为固定技能库的风险暴露认证。
原文摘要 · Abstract (English)
Agent skill libraries are becoming routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and execution assumptions to an agent. This makes retrieval failures more specific than broad irrelevance. A system can find the right capability family yet expose the wrong same-capability representative. We study this failure as same-capability risk-exposure retrieval. Each benchmark unit pairs a helpful skill with a query-specific risky sibling that shares the capability family but differs on an execution-controlling contract, such as the required resource, precondition, procedure, or artifact. We introduce SameCapRisk-Bench, an auditable benchmark with 1,190 skill-risk units and 1,686 evaluation query cases: 694 marked-sibling units under public library pressure and 496 hard role-flip units where the same two skills swap helpful/risky roles across paired queries. The release records admission evidence, cue/leakage checks, source hashes, family relations, and fixed candidate pools. The benchmark reports helpful ranking together with harmful sibling rate (HSR@K), the top-K exposure of the marked risky sibling. On this benchmark, public SkillRouter, SkillRet, and R3-Skill retrieve helpful skills at high Recall@3 (0.848--0.888) but also expose marked risky siblings frequently (HSR@3 0.346--0.372). A fully public score-and-cluster pipeline lowers HSR@3 to 0.128--0.182, with Recall@3 of 0.713--0.776. Under a benchmark-trained reference scorer, public text-cluster and controlled resolvers reach HSR@3 0.012 and 0.007; the latter attains Recall@3 0.833. Skill retrieval should therefore report both capability matching and same-family risk exposure, with HSR serving as a targeted exposure certificate for fixed skill libraries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。