arXiv:2608.12444cs.CRcs.LG2026-08

提出可避免空洞结果的自动化安全决策风险认证方法

Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT\&CK-Aligned Triage as a Worked Instance

  • 构建决策契约理论,揭示错误在自动化、人工延迟和语义遮蔽间的守恒关系
  • 在3个入侵检测数据集上,90.3%配置下误归因风险不超目标值,平均正确自动化率达83.4%
  • 适用于大模型安全告警研判,能区分可修复的阈值偏差与本质能力不足

自动化决策的风险上限可能通过完全不行动实现,导致证书失效。我们证明这种现象具有结构性:任何风险证书都依赖于决策契约,即系统处理的输入及其输出正确的语义关系;削弱任一要素都会隐藏基础分类器的误差。本文提出决策契约理论:包含错误守恒定律,说明错误仅在有害自动化、人工延迟和语义遮蔽间转移;无标签单例容量诊断可识别结构性无能;风险可行的精细化修正能分离可恢复的阈值错配与受风险约束的无能。以基于大模型的入侵检测中对ATT&CK对齐告警的分诊为实例,该设置曾暴露空洞失败问题。在3个IDS数据集、6个大模型和4种误报率阈值下,实证显示90.3%配置中误归因风险不超过目标值,平均正确自动化率达83.4%。容量诊断解释了所有低效配置;其细化版本将真实错配与风险受限无能分离,经替代阈值验证;训练稳定性重跑未发现结构性无能实例;真实细粒度攻击子类型标签证实粗粒度映射下的共性传递关系,存在小但非零的遮蔽质量。

原文摘要 · Abstract (English)

An unconditional risk bound on automated decisions can be satisfied without automating anything, since a selector that never acts drives the bound to zero. We show this is structural: any risk certificate is defined over a decision contract, the inputs a system acts on plus the semantic relation under which an output counts correct, and weakening either hides base-classifier error. We develop a decision-contract theory: an error-conservation law showing error is only reassigned among harmful automation, human deferral, and semantic masking; a label-free singleton capacity certifying structural incapacity, with a risk-feasible refinement separating recoverable threshold misalignment from risk-constrained incapacity; and a non-degenerate actionability certificate excluding all-abstain solutions by construction. We instantiate this on ATT\&CK-aligned alert triage for LLM-based intrusion detection, the setting that exposed the vacuity failure. Across 3 IDS datasets, 6 LLMs, and 4 error-rate thresholds, empirical false-attribution risk stays at or below target in 90.3% of configurations, with 83.4% mean correct automation. The capacity diagnostic explains every low-utility configuration; its refinement separates genuine misalignment from risk-constrained incapacity, confirmed by an exhibited alternative threshold; a training-stability re-run finds no confirmed structural-incapacity instance; and real fine-grained attack-subtype labels confirm the coarsening-transfer identity under a genuine many-to-one map, with small but non-zero masking mass.

安全决策风险认证大模型安全入侵检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。