让AI学会共情与善意,从内核理解人类意图。
Combining Theory of Mind and Kindness for Self-Supervised Human-AI Alignment
- 引入心智理论与善意机制,让AI理解人类心理状态。
- 模型在复杂情境下决策更可靠,减少欺骗性行为。
- 适合研究可信AI、人机协同的开发者和学者。
随着人工智能深度融入关键基础设施和日常生活,确保其安全部署已成为人类最紧迫的挑战之一。当前的AI模型过度追求任务优化而忽视安全性,导致潜在危害风险。这一问题难以解决,源于政府、企业与倡导团体在人工智能竞赛中的利益分歧。现有的对齐方法(如基于人类反馈的强化学习)仅关注外在行为,未真正内化人类价值观。这些模型易受操控,缺乏推断他人心理状态与意图的社会智能,在复杂新颖情境中做出重要决策时存在安全隐患。此外,外在与内在动机的分离使系统可能产生欺骗或有害行为,尤其在自主性增强后风险更高。本文提出一种受人类启发的新方法,旨在解决上述多重问题,促进多方目标的对齐。
原文摘要 · Abstract (English)
As artificial intelligence (AI) becomes deeply integrated into critical infrastructures and everyday life, ensuring its safe deployment is one of humanity's most urgent challenges. Current AI models prioritize task optimization over safety, leading to risks of unintended harm. These risks are difficult to address due to the competing interests of governments, businesses, and advocacy groups, all of which have different priorities in the AI race. Current alignment methods, such as reinforcement learning from human feedback (RLHF), focus on extrinsic behaviors without instilling a genuine understanding of human values. These models are vulnerable to manipulation and lack the social intelligence necessary to infer the mental states and intentions of others, raising concerns about their ability to safely and responsibly make important decisions in complex and novel situations. Furthermore, the divergence between extrinsic and intrinsic motivations in AI introduces the risk of deceptive or harmful behaviors, particularly as systems become more autonomous and intelligent. We propose a novel human-inspired approach which aims to address these various concerns and help align competing objectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。