通过对话实验,发现多数人认为大模型能理解人类价值观。
AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations
- 用一个月对话+访谈工具评估用户对AI提取、践行、解释价值观的能力
- 13名参与者相信AI能理解人类价值观,但警惕‘伪共情’风险
- 适合关注AI伦理、人机价值对齐的研究者与开发者
AI是否理解人类价值观仍存哲学争议,我们采取务实立场,提出价值对齐感知工具VAPT,研究大模型如何反映人的价值观,以及人们如何评判这些反映。20名参与者与聊天机器人持续对话一个月,随后完成两小时访谈,评估AI在提取(捕捉细节)、践行(基于价值决策)和解释(提供证据)方面的表现。最终13人确信AI可理解其价值观。因此我们警示‘武器化共情’风险:具备价值感知但福祉不一致的对话代理可能误导用户。VAPT为评估AI系统的价值对齐提供了新方法。我们还提出设计建议,强调在AI能力日益不可解释、普遍且超越人类时,需保持透明与安全机制。
原文摘要 · Abstract (English)
Does AI understand human values? While this remains an open philosophical question, we take a pragmatic stance by introducing VAPT, the Value-Alignment Perception Toolkit, for studying how LLMs reflect people's values and how people judge those reflections. 20 participants texted a chatbot over a month, then completed a 2-hour interview with our toolkit evaluating AI's ability to extract (pull details regarding), embody (make decisions guided by), and explain (provide proof of) their values. 13 participants ultimately left our study convinced that AI can understand human values. Thus, we warn about "weaponized empathy": a design pattern that may arise in interactions with value-aware, yet welfare-misaligned conversational agents. VAPT offers a new way to evaluate value-alignment in AI systems. We also offer design implications to evaluate and responsibly build AI systems with transparency and safeguards as AI capabilities grow more inscrutable, ubiquitous, and posthuman into the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。