arXiv:2507.11473cs.AIcs.LG2025-07被引 217

通过监控AI的思维链,可发现潜在恶意行为。

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

  • 在AI生成人类语言思维过程时进行实时监控。
  • 当前方法仍会漏检部分恶意行为,但具潜在价值。
  • 适合关注前沿AI安全的开发者与研究者。

能够以人类语言进行“思考”的AI系统为人工智能安全提供了独特机会:我们可以通过监测其思维链(Chain of Thought, CoT)来识别潜在的恶意意图。尽管像所有已知的AI监督方法一样,CoT监控并不完善,仍可能遗漏某些不当行为,但其展现出的潜力值得进一步研究。本文建议将CoT监控作为现有安全手段的补充,并投入更多资源。由于CoT监控可能具有脆弱性,建议前沿模型开发者在研发过程中评估其对CoT监控能力的影响。

原文摘要 · Abstract (English)

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.

AI安全思维链监控机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。