arXiv:2512.15722cs.CYcs.AI2025-12被引 3

用大模型识别文本中的人类价值观,让机器决策更符合人性。

Value Lens: Using Large Language Models to Understand Human Values

  • 分两阶段:先构建价值理论,再检测文本中的价值
  • 大模型自动识别并由另一模型审核,准确率超同类方法
  • 适合伦理对齐、AI决策透明化等研究者使用

自主决策系统日益广泛应用于计算机系统,其决策必须与人类价值观保持一致。为此,系统需评估自身选择是否体现人类价值观。本文提出Value Lens,一种基于生成式人工智能(特别是大型语言模型)的文本型价值检测模型。该模型分两阶段运行:第一阶段由大模型生成基于既有价值理论的描述,经专家验证;第二阶段采用一对大模型,一个负责检测价值存在,另一个作为评审进行复核。实验表明,Value Lens在性能上可媲美甚至优于其他同类方法。

原文摘要 · Abstract (English)

The autonomous decision-making process, which is increasingly applied to computer systems, requires that the choices made by these systems align with human values. In this context, systems must assess how well their decisions reflect human values. To achieve this, it is essential to identify whether each available action promotes or undermines these values. This article presents Value Lens, a text-based model designed to detect human values using generative artificial intelligence, specifically Large Language Models (LLMs). The proposed model operates in two stages: the first aims to formulate a formal theory of values, while the second focuses on identifying these values within a given text. In the first stage, an LLM generates a description based on the established theory of values, which experts then verify. In the second stage, a pair of LLMs is employed: one LLM detects the presence of values, and the second acts as a critic and reviewer of the detection process. The results indicate that Value Lens performs comparably to, and even exceeds, the effectiveness of other models that apply different methods for similar tasks.

大模型价值观伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。