让大模型写注释,帮小模型更好识别恶意代码。
Basic Legibility Protocols Improve Trusted Monitoring
- 让不可信模型主动添加详细注释,提升可解释性
- 在编码任务中,安全率提升且不降低性能
- 适合需要可信监督的AI部署场景
AI控制研究旨在开发控制协议:防止不可信AI系统在部署中采取有害行为的安全技术。由于人工监督成本高,一种方法是信任监控,即由较弱的可信模型监督较强的不可信模型——但当不可信模型的行为超出监控者的理解能力时,此方法常失效。我们提出可读性协议,促使不可信模型采取更易被监控者评估的动作。在APPS编码环境中,我们测试了允许不可信模型详尽添加代码注释的协议,不同于以往为防欺骗而删除注释的做法。结果表明:(i) 注释协议在不牺牲任务表现的前提下提升了安全性;(ii) 注释对诚实代码的收益更大,因其通常有自然解释以消除监控疑虑,而后门代码往往缺乏合理理由;(iii) 监控模型越强,注释带来的安全增益越大,因更强模型能更好区分真实解释与表面合理但虚假的说辞。
原文摘要 · Abstract (English)
The AI Control research agenda aims to develop control protocols: safety techniques that prevent untrusted AI systems from taking harmful actions during deployment. Because human oversight is expensive, one approach is trusted monitoring, where weaker, trusted models oversee stronger, untrusted models$\unicode{x2013}$but this often fails when the untrusted model's actions exceed the monitor's comprehension. We introduce legibility protocols, which encourage the untrusted model to take actions that are easier for a monitor to evaluate. We perform control evaluations in the APPS coding setting, where an adversarial agent attempts to write backdoored code without detection. We study legibility protocols that allow the untrusted model to thoroughly document its code with comments$\unicode{x2013}$in contrast to prior work, which removed comments to prevent deceptive ones. We find that: (i) commenting protocols improve safety without sacrificing task performance relative to comment-removal baselines; (ii) commenting disproportionately benefits honest code, which typically has a natural explanation that resolves monitor suspicion, whereas backdoored code frequently lacks an easy justification; (iii) gains from commenting increase with monitor strength, as stronger monitors better distinguish genuine justifications from only superficially plausible ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。