arXiv:2502.02444cs.CLcs.AI2025-02ACL被引 7

用心理语言学方法构建大模型价值观体系,提升安全与对齐效果

Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models

  • 基于心理词汇学原理生成可扩展的大模型价值体系
  • 新体系在心理标准和安全预测上优于经典舒瓦茨价值观
  • 适合研究大模型对齐、伦理评估的AI安全方向学者

价值观是个人与集体认知、感知和行为的核心驱动力。如舒瓦茨基本人类价值观理论所定义的价值体系,揭示了价值间的层级关系与相互作用,为决策与社会动态的跨学科研究提供基础。近年来,大语言模型(LLMs)内在价值观的模糊性引发关注。尽管已有大量研究致力于评估、理解与对齐模型价值观,但基于心理学的系统性价值体系仍属空白。本文提出生成式心理词汇法(GPLA),一种可扩展、可适配且理论驱动的方法,用于构建大模型价值体系。基于GPLA,我们设计了一个符合心理学依据的五因素价值体系。通过三个结合心理原则与前沿AI目标的基准任务进行系统验证,结果表明:该体系满足标准心理标准,更准确捕捉大模型价值观,显著提升安全预测能力,并增强模型对齐效果,优于经典的舒瓦茨价值观体系。

原文摘要 · Abstract (English)

Values are core drivers of individual and collective perception, cognition, and behavior. Value systems, such as Schwartz's Theory of Basic Human Values, delineate the hierarchy and interplay among these values, enabling cross-disciplinary investigations into decision-making and societal dynamics. Recently, the rise of Large Language Models (LLMs) has raised concerns regarding their elusive intrinsic values. Despite growing efforts in evaluating, understanding, and aligning LLM values, a psychologically grounded LLM value system remains underexplored. This study addresses the gap by introducing the Generative Psycho-Lexical Approach (GPLA), a scalable, adaptable, and theoretically informed method for constructing value systems. Leveraging GPLA, we propose a psychologically grounded five-factor value system tailored for LLMs. For systematic validation, we present three benchmarking tasks that integrate psychological principles with cutting-edge AI priorities. Our results reveal that the proposed value system meets standard psychological criteria, better captures LLM values, improves LLM safety prediction, and enhances LLM alignment, when compared to the canonical Schwartz's values.

大模型对齐价值观建模心理语言学安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。