arXiv:2510.19476cs.LGcs.AI2025-10

通过思维链监控构建可信安全案例,防范高危AI风险。

A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring

  • 用思维链监控检测模型是否具备危险能力
  • 提出双层安全验证:无思维链时无危险能力,有则可被监控
  • 针对隐蔽推理形式提出三类威胁分析,适配安全研究者

随着人工智能系统接近具有潜在危险能力的临界点,传统安全论证已不足以保障安全。本文提出基于思维链(CoT)监控构建安全案例的路线图,并明确研究方向。我们认为,思维链监控可支持控制性与可信性两类安全案例。安全案例分为两部分:(1) 确保模型在无思维链状态下不具备危险能力;(2) 保证任何由思维链激活的危险能力均可被思维链监控识别。系统分析了两种影响可监控性的威胁:神经语言(neuralese)和编码推理,将其归类为三种形式(语言漂移、隐写术、异质推理),并探讨其潜在驱动因素。评估现有及新提出的保持思维链忠实性的技术。对于产生不可监控推理的场景,探索从非可监控思维链中提取可监控思维链的可能性。为评估思维链监控安全案例的可行性,建立预测市场以聚合对关键技术里程碑的共识预测。

原文摘要 · Abstract (English)

As AI systems approach dangerous capability levels where inability safety cases become insufficient, we need alternative approaches to ensure safety. This paper presents a roadmap for constructing safety cases based on chain-of-thought (CoT) monitoring in reasoning models and outlines our research agenda. We argue that CoT monitoring might support both control and trustworthiness safety cases. We propose a two-part safety case: (1) establishing that models lack dangerous capabilities when operating without their CoT, and (2) ensuring that any dangerous capabilities enabled by a CoT are detectable by CoT monitoring. We systematically examine two threats to monitorability: neuralese and encoded reasoning, which we categorize into three forms (linguistic drift, steganography, and alien reasoning) and analyze their potential drivers. We evaluate existing and novel techniques for maintaining CoT faithfulness. For cases where models produce non-monitorable reasoning, we explore the possibility of extracting a monitorable CoT from a non-monitorable CoT. To assess the viability of CoT monitoring safety cases, we establish prediction markets to aggregate forecasts on key technical milestones influencing their feasibility.

AI安全思维链监控机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。