arXiv:2605.16198cs.AIcs.CY2026-05被引 3

用形式化方法监控大模型行为,确保其符合安全与合规要求。

Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems

论文配图:Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
图 1 · 摘自论文原文
  • 结合线性时序逻辑(LTL)与大模型,实现对复杂行为约束的审计与实时监测。
  • 小模型标注器在检测违规上媲美甚至超越前沿大模型判官,且显著降低违规率。
  • 适用于大模型开发者、监管机构及第三方评估者,提升AI系统可问责性。

本文研究AI治理中的关键问题:如何在人工智能开发全生命周期中监控和审计AI产品与服务,涵盖部署前测试到部署后审计。通过融合形式化方法与当前最先进机器学习技术,提出一套方法,使开发者及第三方评估者能够对黑箱大模型(如大语言模型)进行离线审计与在线运行时监控,以检测安全、规范、规则等随时间演化的行为约束是否被违反。我们进一步提供基于采样的预测性监控技术,并引入运行时干预监控机制,提前预防并缓解潜在违规。实验表明,利用线性时序逻辑(LTL)的形式语法与语义,本方法在检测时序行为约束违规方面优于大模型基线;即使是小型模型标注器,也能达到或超过前沿大模型判官的表现。预测与干预监控显著降低了大模型代理的违规率,同时基本保持任务性能。控制实验还显示,大模型在事件距离增加、约束数量增多或命题数上升时,其时序推理准确率明显下降。

原文摘要 · Abstract (English)

We examine one particular dimension of AI governance: how to monitor and audit AI-enabled products and services throughout the AI development lifecycle, from pre-deployment testing to post-deployment auditing. Combining principles from formal methods with SoTA machine learning, we propose techniques that enable AI-enabled product and service developers, as well as third party AI developers and evaluators, to perform offline auditing and online (runtime) monitoring of product-specific (temporally extended) behavioral constraints such as safety constraints, norms, rules and regulations with respect to black-box advanced AI systems, notably LLMs. We further provide practical techniques for predictive monitoring, such as sampling-based methods, and we introduce intervening monitors that act at runtime to preempt and potentially mitigate predicted violations. Experimental results show that by exploiting the formal syntax and semantics of Linear Temporal Logic (LTL), our proposed auditing and monitoring techniques are superior to LLM baseline methods in detecting violations of temporally extended behavioral constraints; with our approach, even small-model labelers match or exceed frontier LLM judges. Our predictive and intervening monitors significantly reduce the violation rates of LLM-based agents while largely preserving task performance. We further show through controlled experiments that LLMs' temporal reasoning shows a pronounced degradation in accuracy with increasing event distance, number of constraints, and number of propositions.

大模型治理形式化验证行为监控合规审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。