arXiv:2508.07673cs.AIcs.LG2025-08AAAI

将自动代理决策转化为可度量向量,实现与人类价值观对齐。

Ethics2vec: aligning automatic agents and human preferences

  • 提出Ethics2Vec方法,将代理行为映射为多维向量
  • 可在不相容价值间建立可比较的度量空间
  • 适用于医疗、自动驾驶等高伦理敏感场景

尽管智能代理旨在提升人类体验或效率,但从人类视角看,很难理解代理行为中隐含或明确嵌入的伦理价值。这就是著名的对齐问题——设计与人类价值观、目标和偏好一致的AI系统。该问题尤为棘手,因为多数人类伦理考量涉及不可比(即不可测量且/或不可比较)的价值与标准。例如,一个为癌症患者开药的医疗代理,如何权衡人类生命价值与治疗成本?只有在定义了可度量的共同空间时,人类与人工价值之间的对齐才可能实现。本文提出将传统的Anything2vec方法扩展至伦理领域,该方法已在自然语言处理、推荐系统和图分析等难以量化领域取得成功。本文提出一种将自动代理决策(或控制律)策略映射为多变量向量表示的方法,可用于比较和评估其与人类价值观的对齐程度。首先在二元决策场景下介绍Ethics2Vec,随后讨论自驾车等自动控制场景中的控制律向量化,展示该方法的可扩展性。

原文摘要 · Abstract (English)

Though intelligent agents are supposed to improve human experience (or make it more efficient), it is hard from a human perspective to grasp the ethical values which are explicitly or implicitly embedded in an agent behaviour. This is the well-known problem of alignment, which refers to the challenge of designing AI systems that align with human values, goals and preferences. This problem is particularly challenging since most human ethical considerations refer to \emph{incommensurable} (i.e. non-measurable and/or incomparable) values and criteria. Consider, for instance, a medical agent prescribing a treatment to a cancerous patient. How could it take into account (and/or weigh) incommensurable aspects like the value of a human life and the cost of the treatment? Now, the alignment between human and artificial values is possible only if we define a common space where a metric can be defined and used. This paper proposes to extend to ethics the conventional Anything2vec approach, which has been successful in plenty of similar and hard-to-quantify domains (ranging from natural language processing to recommendation systems and graph analysis). This paper proposes a way to map an automatic agent decision-making (or control law) strategy to a multivariate vector representation, which can be used to compare and assess the alignment with human values. The Ethics2Vec method is first introduced in the case of an automatic agent performing binary decision-making. Then, a vectorisation of an automatic control law (like in the case of a self-driving car) is discussed to show how the approach can be extended to automatic control settings.

伦理对齐向量表示自动决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。