提出评估人类与大模型价值对齐的框架,发现两者在国家安全等议题上存在显著分歧。
ValueCompass: A Framework for Measuring Contextual Value Alignment Between Human and LLMs
- 基于心理学理论构建价值观框架,系统评估人与AI对齐程度
- 实测显示人类重视'国家安全'等价值,而大模型多拒绝此类主张
- 不同场景下价值观差异明显,提示需发展情境感知的对齐策略
随着人工智能系统日益复杂,确保其与多元个体及社会价值保持一致变得愈发重要。如何捕捉基本人类价值观并评估人工智能系统的对齐程度?本文提出ValueCompass框架,该框架基于心理学理论和系统综述,用于识别与评估人类-人工智能对齐情况。我们在四个真实场景——协作写作、教育、公共部门和医疗健康——中应用该框架,测量人类与大型语言模型(LLMs)的价值对齐度。研究发现,人类频繁支持如‘国家安全部’等价值观,而这些观点被大多数大模型拒绝。此外,不同场景中的价值观表现存在差异,凸显了发展情境感知型人工智能对齐策略的必要性。本工作为理解人机对齐的设计空间提供了关键洞见,奠定了负责任反映社会价值与伦理的人工智能系统的基础。
原文摘要 · Abstract (English)
As AI systems become more advanced, ensuring their alignment with a diverse range of individuals and societal values becomes increasingly critical. But how can we capture fundamental human values and assess the degree to which AI systems align with them? We introduce ValueCompass, a framework of fundamental values, grounded in psychological theory and a systematic review, to identify and evaluate human-AI alignment. We apply ValueCompass to measure the value alignment of humans and large language models (LLMs) across four real-world scenarios: collaborative writing, education, public sectors, and healthcare. Our findings reveal concerning misalignments between humans and LLMs, such as humans frequently endorse values like "National Security" which were largely rejected by LLMs. We also observe that values differ across scenarios, highlighting the need for context-aware AI alignment strategies. This work provides valuable insights into the design space of human-AI alignment, laying the foundations for developing AI systems that responsibly reflect societal values and ethics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。