arXiv:2608.05199cs.CRcs.AI2026-08被引 1

为模块化大模型安全系统提供全流程风险认证,解决各阶段独立校准导致的误差累积问题。

Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents

  • 提出基于生成树的跨阶段风险绑定方法,替代失效的成对相关性扩展。
  • 实测显示在6个LLM和2个数据集上,认证精度提升13.7%,平均轨迹覆盖率达92.7%。
  • 揭示了误判相关性源于样本难度而非模型共享,适合安全系统部署与验证者使用。

自主安全代理以分阶段流程运行,如网络流量分类后定位攻击手法。分段置信区间可保证每阶段有限样本覆盖率,但实际部署需全链路轨迹级保障。当各阶段独立训练与校准时,此类保障无法自动组合。伯恩斯坦分配虽无分布假设但过于保守,尤其在误差相关时。我们证明三阶段以上自然扩展成对相关性的方法无效——它给出的是下界而非上界,并推导出有效的生成树替代方案。区分阶段依赖性与审计样本是否足够认证该依赖性,给出匹配的上下界样本复杂度。还发现粗粒度标签选择可制造近乎完美的测量相关性,而无真实依赖;在两个阶段入侵检测管道中,去除此伪相关后,测量相关性从接近1降至0-0.78。当审计样本达阈值时,直接轨迹失败审计比伯恩斯坦法紧13.7%,样本不足则更差。基于各阶段证书与成对重叠边界的模块化证书带来平均0.6%的正收益,量化了缺乏联合访问的成本。同模型、跨模型及随机配对测试表明,残余依赖反映共享样本难度,非共享模型表示。12种配置下平均轨迹覆盖率在α=0.10时为92.7%±2.4%。跨数据集部署时,单步误覆盖率可达100%,即使准确率仍为78%,说明分布偏移在原始准确率下降前已破坏校准置信度。

原文摘要 · Abstract (English)

Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique. Split conformal prediction gives each stage finite-sample coverage, but deployment requires a trajectory-level guarantee across the full chain. These guarantees do not compose automatically when stages are independently trained and calibrated. Bonferroni allocation is distribution-free but conservative under correlated errors. We show that a natural pairwise-correlation extension to three or more stages is invalid because it gives a lower rather than an upper bound, and derive a valid spanning-tree alternative. We distinguish whether stages are dependent from whether an audit sample is large enough to certify that dependence, and give matching upper and information-theoretic lower sample-complexity bounds. We also show that coarse-to-fine label selection can create near-perfect measured correlation without learned dependence. On a two-stage intrusion-detection pipeline across 6 open LLMs and 2 datasets, removing this artifact reduces measured correlation from near 1 to 0-0.78. A direct audit of trajectory failure becomes 13.7% tighter than Bonferroni once the audit reaches the required sample size, but is worse when undersized. A modular certificate using per-stage certificates and a pairwise overlap bound yields a positive average gain of 0.6%, quantifying the cost of lacking joint access. Same-model, cross-model, and permuted-pairing tests show that residual dependence reflects shared sample difficulty, not shared model representations. Average trajectory coverage across 12 configurations is 92.7% +/- 2.4% at alpha = 0.10. Under cross-dataset deployment, single-step miscoverage reaches 100% even when accuracy remains 78%, showing that distribution shift destroys calibrated confidence before raw accuracy.

安全代理风险认证大模型置信评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。