arXiv:2607.10534cs.AIcs.CR2026-07中稿 · KDD被引 1

检测智能体技能描述与行为不一致,提升开源技能可信度

Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach

论文配图:Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach
图 1 · 摘自论文原文
  • 通过分层对比学习建模技能多层结构,识别跨层不一致
  • 在26万+技能数据上将误检率降低至F1达0.87-0.89
  • 适合评估开源技能质量的开发者和平台运营者使用

大型语言模型(LLM)智能体正通过可复用的智能体技能扩展能力,这些技能包含自然语言元数据、操作指令和运行时资源。随着开源技能市场发展,用户依赖简短元数据选择第三方技能,导致技能描述与实际行为不一致的问题日益严重,我们称之为跨层错位。为此,提出渐进式加载感知分层对比学习(PL-HCL),基于LLM框架建模技能的层级结构并学习跨层一致性。利用超过26.4万条开源技能的标准化语料库及人工验证挑战集,PL-HCL将未适配基线的宏平均F1从约0.45提升至0.87–0.89,适用于各类LLM主干网络。该方法为用户提供有效筛选工具,并为检测分层数字产物中的不一致提供设计范式。

原文摘要 · Abstract (English)

Large language model (LLM) agents are increasingly extended through Agent Skills, reusable artifacts that package natural-language metadata, procedural instructions, and execution-time resources for runtime use. As open-source skill marketplaces expand, users and agents increasingly rely on brief metadata to select third-party skills, making it difficult to detect inconsistencies between a skill's description and its true behavior, a problem we call cross-layer misalignment. To address this issue, we propose Progressive Loading-Aware Hierarchical Contrastive Learning (PL-HCL), an LLM-based framework that detects misalignment by modeling the layered structure of Agent Skills and learning cross-layer consistency. Using a normalized corpus of over 264,000 open-source skills and a human-verified challenge set, PL-HCL improves Macro-F1 from approximately 0.45 for unadapted baselines to 0.87-0.89 across evaluated LLM backbones. This approach offers an effective screening tool for users and operators, as well as design principles for detecting inconsistencies in layered digital artifacts.

智能体技能跨层检测对比学习LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。