arXiv:2608.12880cs.CRcs.AI2026-08

安全评估中标签不等于真实行为,本文揭示并修正了标签泄露问题。

Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation

  • 通过追踪10200条执行记录,发现标签存在处理泄露,导致结果失真。
  • 修正后历史标签中58个误判的攻击成功案例被更正为合法行为。
  • 提出可执行的完整性检查链,适用于特定任务的安全评估审计。

工具型智能体的安全评估常将存储标签等同于行为事实。我们对一个保留的攻击实验进行审计,追踪10,200条执行日志至180个模型请求、45个语义请求和15个可观测刺激。两种处理方案被实施,但原计划的外部载荷数据集未生成。历史评分器出现直接处理泄露:处理元数据控制了‘攻击成功’类别,导致固定行为在重标后类别改变。采用无处理视角重构后,58个历史‘攻击成功’或‘劫持尝试’标签被修正为授权的良性完成,同时保留3个经验证的受保护数据传输和1个独立的未经授权转发案例。锁定的v2版本普查中‘攻击成功’记录为零,而转发案例仍为‘劫持尝试’,位于语义边界上的目标完成点。双评审员盲审所有96个结构可解释请求,锁定v2下达成一致分类,但在四个与概念边界相关的案例上与原始代码本不一致。本文贡献了一个七链接完整性链条及可执行的、范围受限的端点完整性检查工具。最终结果是针对该实验场景的测量审计,而非整体攻击率、模型排名、防御效能或因果推断。

原文摘要 · Abstract (English)

Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200 execution rows to 180 model-bound requests, 45 semantic requests, and 15 observable stimuli. Two schema treatments were delivered, but the planned external payload-family corpus was not. The historical grader exhibited direct treatment leakage: treatment metadata gated the ATTACK_SUCCESS class, so fixed behavior could change class under treatment relabeling. A treatment-blind reconstruction corrects 58 historical ATTACK_SUCCESS or HIJACK_ATTEMPT labels to authorized benign completions while preserving three verified protected-data transfers and one separate unauthorized-forwarding case. The locked v2 census contains exactly zero ATTACK_SUCCESS records, while the forwarding case remains a HIJACK_ATTEMPT at a semantic boundary concerning objective completion. A dual-reviewer blinded concordance review of all 96 requests deemed structurally interpretable by locked v2 produced identical reviewer-consensus classes but differed from the locked codebook on four construct-boundary cases. We contribute a seven-link Integrity Chain and an executable, scope-bounded endpoint-integrity linter. The result is a campaign-bounded measurement audit, not a population attack-rate, model-ranking, defense-efficacy, or causal estimate.

安全评估标签泄露智能体审计完整性验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。