用自洽方法检验AI安全分析工具自身,防止其产生不可靠的安全结论。
Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA
- 将STPA方法应用于AI分析工具本身,构建自我验证的治理框架。
- 从工具设计中自动推导出21条工具原则和8条元安全原则,覆盖率达100%。
- 适用于需要高可信安全分析的自动驾驶、医疗系统等关键领域。
大型语言模型(LLMs)正被广泛用于生成安全分析中的风险、危害、不安全控制动作(UCAs)及安全约束等成果,尤其在系统理论过程分析(STPA)流程中。然而现有研究存在盲区:所有系统都被分析,唯独负责分析的LLM工具本身未被审视——它可能虚构标准、生成无法验证的约束,且缺乏从提示到输出的审计痕迹。本文严肃回应‘谁来分析分析者?’这一问题,提出宪法式元STPA(Constitutional Meta-STPA),通过闭环机制对一类AI辅助安全工具进行元级STPA,从损失→危害→不安全控制动作→约束链中推导而非声明其治理宪章,形成21条工具原则与8条元安全原则,并绑定代码执行点。我们定义了基于原则集 $P$(|P|=29)的宪章边际覆盖率算子,给出完备性引理以分离覆盖率与模型和扫描器的影响。实验发现:前沿模型组合(claude-opus-4.8 + claude-sonnet-4)可恢复全部21条规范原则与8条治理原则;较弱组合恢复12/21与3/8;同一套8条治理原则亦在另一独立开发工具中复现,表明元层受限于模型能力,非宪章本身。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs), and safety constraints, inside rigorous processes such as Systems-Theoretic Process Analysis (STPA). Yet a blind spot runs through this fast-growing literature: every system gets analysed except the LLM-assisted tool doing the analysing, which is itself a safety-relevant system that can hallucinate standards, emit unverifiable constraints, and leave no audit trail from prompt to artifact. We take seriously the question the field has skipped -- {who analyses the analyser?} and answer it by turning STPA on the tool itself. We present \{Constitutional Meta-STPA}, an LLM-assisted STPA tool built around a closed loop: the tool runs a {meta-STPA} of the class of AI-assisted safety tools and {derives} rather than asserts, its governance constitution from the resulting loss$\to$hazard$\to$UCA$\to$constraint chain, yielding a published constitution of $21$ Tool Principles and $8$ Meta-Safety Principles, each bound to a code enforcement point. We formalise the measured object as a constitution-marginal coverage operator over a principle set $P$ ($|P|{=}29$) with a soundness lemma that isolates coverage from model and scanner, and report four findings. {(i)~Self-derivation:} a frontier ensemble ({claude-opus-4.8}${+}${claude-sonnet-4}) recovers $18/21$ canonical and all $8/8$ governance principles from the tool's own design, while a weaker pair recovers $12/21$ and $3/8$, so the meta layer is model-limited, not constitution-limited, and the same $8/8$ re-emerge from a second, independently authored tool.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。