评估大模型推理过程的可监控性,发现更高推理成本能提升安全性。
Monitoring Monitorability
- 设计三类评估方法与新指标,量化模型推理过程是否易被监控
- 长推理链更易监控,强化学习训练不降低监控效果
- 通过追问和提供追问推理链,可显著提升监控能力
现代AI系统决策过程的可观测性对安全部署至关重要。尽管当前推理模型的思维链(CoT)监控已有效识别异常行为,但其可监控性可能在不同训练方式、数据来源或模型规模下变得脆弱。为此,本文提出三种评估范式(干预、过程、结果属性)及新的可监控性度量,并构建全面评估体系。实验表明,该方法能识别出刻意隐藏思维链的简化模型,且相比仅监控动作,监控思维链在实际中更有效。对比多个前沿模型发现,多数模型具备一定可监控性但非完美。研究还发现,推理时间计算资源增加、强化学习优化和预训练模型规模扩大均影响可监控性:更长的思维链普遍更易监控,而强化学习在当前前沿规模下未显著降低可监控性。值得注意的是,在低推理强度下使用小模型以高推理强度达到同等能力,可获得更高可监控性,代价是更高推理计算开销。进一步分析代理-监控器扩展趋势发现,弱监控器在监控强代理时,若提升其测试时计算资源,可增强可监控性;若赋予其访问思维链,不仅提升性能,还使监控能力随计算投入呈更陡峭上升趋势。最后,通过让模型回答后续问题并提供其后续思维链,可有效提升监控能力。
原文摘要 · Abstract (English)
Observability into the decision making of modern AI systems may be required to safely deploy increasingly capable agents. Monitoring the chain-of-thought (CoT) of today's reasoning models has proven effective for detecting misbehavior. However, this "monitorability" may be fragile under different training procedures, data sources, or even continued system scaling. To measure and track monitorability, we propose three evaluation archetypes (intervention, process, and outcome-property) and a new monitorability metric, and introduce a broad evaluation suite. We demonstrate that these evaluations can catch simple model organisms trained to have obfuscated CoTs, and that CoT monitoring is more effective than action-only monitoring in practical settings. We compare the monitorability of various frontier models and find that most models are fairly, but not perfectly, monitorable. We also evaluate how monitorability scales with inference-time compute, reinforcement learning optimization, and pre-training model size. We find that longer CoTs are generally more monitorable and that RL optimization does not materially decrease monitorability even at the current frontier scale. Notably, we find that for a model at a low reasoning effort, we could instead deploy a smaller model at a higher reasoning effort (thereby matching capabilities) and obtain a higher monitorability, albeit at a higher overall inference compute cost. We further investigate agent-monitor scaling trends and find that scaling a weak monitor's test-time compute when monitoring a strong agent increases monitorability. Giving the weak monitor access to CoT not only improves monitorability, but it steepens the monitor's test-time compute to monitorability scaling trend. Finally, we show we can improve monitorability by asking models follow-up questions and giving their follow-up CoT to the monitor.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。