让人类与AI双向适应,提升协作效率与安全。
Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation
- 提出双向认知对齐机制,实现人与AI共同进化。
- 导航任务成功率提升至85.5%,协议收敛速度提高332%。
- 适合追求高效人机协作与安全性的系统设计者。
当前基于强化学习的人工智能对齐遵循单向范式,即人工智能顺应人类偏好,而将人类认知视为固定不变。本文提出转向协同对齐(Co-Alignment),通过双向认知对齐(BiCA)实现人类与AI的相互适应。BiCA采用可学习的交互协议、表示映射及KL预算约束,实现受控的共同演化。在协作导航任务中,BiCA达成85.5%的成功率,优于基线70.3%;互适应性提升230%,协议收敛速度提高332%。涌现的交互协议相比手工设计提升84%,且双向适应意外带来23%的分布外鲁棒性提升。46%的协同增益表明,最优协作发生在人与AI能力的交集而非并集,验证了从单向到协同对齐范式的必要性。
原文摘要 · Abstract (English)
Current AI alignment through RLHF follows a single directional paradigm that AI conforms to human preferences while treating human cognition as fixed. We propose a shift to co-alignment through Bidirectional Cognitive Alignment (BiCA), where humans and AI mutually adapt. BiCA uses learnable protocols, representation mapping, and KL-budget constraints for controlled co-evolution. In collaborative navigation, BiCA achieved 85.5% success versus 70.3% baseline, with 230% better mutual adaptation and 332% better protocol convergence. Emergent protocols outperformed handcrafted ones by 84%, while bidirectional adaptation unexpectedly improved safety (+23% out-of-distribution robustness). The 46% synergy improvement demonstrates optimal collaboration exists at the intersection, not union, of human and AI capabilities, validating the shift from single-directional to co-alignment paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。