通过模拟共情机制,让AI更诚实可靠
Towards Safe and Honest AI Agents with Neural Self-Other Overlap
- 让AI自我与他人表征对齐,提升诚实性
- 780亿参数模型欺骗率从100%降至2.7%
- 适合构建可信AI的科研与工程人员参考
随着人工智能系统在关键决策中作用增强,欺骗性AI已成为信任与安全的重大挑战。本文提出自-他重叠(SOO)微调方法,受认知神经科学共情研究启发,旨在对齐AI模型对自我与他者的表征。在7B、27B和78B参数的大语言模型上实验表明:Mistral-7B-Instruct-v0.2的欺骗响应率从73.6%降至17.2%,未影响通用任务表现;Gemma-2-27b-it与CalmeRys-78B-Orpo-v0.1的欺骗率分别从100%降至9.3%和2.7%,能力小幅下降。强化学习场景中,经SOO训练的智能体欺骗行为显著减少。该方法基于对比式自/他参照观测,具备跨架构泛化潜力。当前应用集中于语言模型与简单强化学习环境,但有望推动更广泛领域中可信赖AI的发展。伦理影响与长期效应仍需进一步研究,但SOO代表了AI安全研究的重要进展。
原文摘要 · Abstract (English)
As AI systems increasingly make critical decisions, deceptive AI poses a significant challenge to trust and safety. We present Self-Other Overlap (SOO) fine-tuning, a promising approach in AI Safety that could substantially improve our ability to build honest artificial intelligence. Inspired by cognitive neuroscience research on empathy, SOO aims to align how AI models represent themselves and others. Our experiments on LLMs with 7B, 27B, and 78B parameters demonstrate SOO's efficacy: deceptive responses of Mistral-7B-Instruct-v0.2 dropped from 73.6% to 17.2% with no observed reduction in general task performance, while in Gemma-2-27b-it and CalmeRys-78B-Orpo-v0.1 deceptive responses were reduced from 100% to 9.3% and 2.7%, respectively, with a small impact on capabilities. In reinforcement learning scenarios, SOO-trained agents showed significantly reduced deceptive behavior. SOO's focus on contrastive self and other-referencing observations offers strong potential for generalization across AI architectures. While current applications focus on language models and simple RL environments, SOO could pave the way for more trustworthy AI in broader domains. Ethical implications and long-term effects warrant further investigation, but SOO represents a significant step forward in AI safety research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。