arXiv:2507.21091cs.CYcs.AI2025-07被引 2

从对话日志中挖掘真实价值冲突,构建可落地的AI伦理对齐框架

The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment

  • 基于实际对话日志自下而上提炼核心价值
  • 发现9个核心价值与32种具体价值错配现象
  • 为就业类AI提供可操作的伦理对齐方案

基于生成式人工智能的对话代理(CAs)常面临与人类价值观对齐的挑战。现有对齐方法多采用自上而下的技术或法律规范,但往往脱离实际应用场景,易导致与用户利益偏离。为此,本文提出一种新型自下而上的价值对齐方法,利用ISO价值工程标准中的价值本体框架。分析了来自欧洲某主要就业服务CA的16,908条对话日志中识别出的593个伦理敏感输出,揭示了9个核心价值及32种负面影响用户的值错配情形。研究结果为CA开发者提供了可执行的伦理改进路径,推动更贴近现实情境的价值对齐。

原文摘要 · Abstract (English)

Conversational agents (CAs) based on generative artificial intelligence frequently face challenges ensuring ethical interactions that align with human values. Current value alignment efforts largely rely on top-down approaches, such as technical guidelines or legal value principles. However, these methods tend to be disconnected from the specific contexts in which CAs operate, potentially leading to misalignment with users interests. To address this challenge, we propose a novel, bottom-up approach to value alignment, utilizing the value ontology of the ISO Value-Based Engineering standard for ethical IT design. We analyse 593 ethically sensitive system outputs identified from 16,908 conversational logs of a major European employment service CA to identify core values and instances of value misalignment within real-world interactions. The results revealed nine core values and 32 different value misalignments that negatively impacted users. Our findings provide actionable insights for CA providers seeking to address ethical challenges and achieve more context-sensitive value alignment.

AI伦理对话系统价值对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。