提出可量化推理可读性与覆盖度的实用方法,助力AI安全监控。
A Pragmatic Way to Measure Chain-of-Thought Monitorability
- 用自动评分提示让大模型评估推理链的可读性和覆盖度。
- 前沿模型在复杂任务上仍保持高可监控性。
- 适合关注AI安全的开发者追踪设计对可监控性的影响。
链式思维(CoT)监控为人工智能安全提供了独特机遇,但训练方式或模型架构的改变可能使其失效。为维持可监控性,我们提出一种实用方法,用于衡量其两个关键成分:可读性(人类能否理解推理过程)和覆盖度(推理链是否包含生成最终输出所需的所有信息)。我们通过一个自动评分提示实现该测量,使任何能力足够的大模型都能计算现有CoT的可读性和覆盖度。经合成退化数据验证后,我们在多个前沿模型及挑战性基准上应用该方法,发现这些模型仍具备较高的可监控性。我们公开了该指标及完整自动评分提示,供开发者追踪设计决策对可监控性的影响。尽管当前提示版本仍处于持续开发中,但我们希望社区能尽早使用。本方法旨在评估CoT的默认可监控性,应作为对抗性压力测试的补充,而非替代,以检验模型对故意规避行为的鲁棒性。
原文摘要 · Abstract (English)
While Chain-of-Thought (CoT) monitoring offers a unique opportunity for AI safety, this opportunity could be lost through shifts in training practices or model architecture. To help preserve monitorability, we propose a pragmatic way to measure two components of it: legibility (whether the reasoning can be followed by a human) and coverage (whether the CoT contains all the reasoning needed for a human to also produce the final output). We implement these metrics with an autorater prompt that enables any capable LLM to compute the legibility and coverage of existing CoTs. After sanity-checking our prompted autorater with synthetic CoT degradations, we apply it to several frontier models on challenging benchmarks, finding that they exhibit high monitorability. We present these metrics, including our complete autorater prompt, as a tool for developers to track how design decisions impact monitorability. While the exact prompt we share is still a preliminary version under ongoing development, we are sharing it now in the hopes that others in the community will find it useful. Our method helps measure the default monitorability of CoT - it should be seen as a complement, not a replacement, for the adversarial stress-testing needed to test robustness against deliberately evasive models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。