arXiv:2606.06674cs.CLcs.CY2026-06中稿 · the 2026 ACM Confe…

发现人们对AI的期待差异巨大,单一奖励模型难以满足真实多元需求。

What Do People Actually Want From AI? Mapping Preference Plurality

  • 分析1500份跨75国问卷,揭示用户对AI需求高度分化
  • 仅49%人要求真实性,且定义各异:有溯源、专家意见或反主流观点
  • 现有二元比较方法忽略上下文差异,无法捕捉复杂偏好

大型语言模型常通过人类反馈强化学习(RLHF)对齐人类偏好,但该方法存在聚合冲突需求、样本不具代表性及仅用二元对比等缺陷。基于来自75个国家的PRISM数据集中1500份开放式回应,本文探究人们真正希望从AI获得什么,并揭示当前对齐方法的实质问题。结果显示,多数价值诉求仅被少于四分之一受访者提出,唯有‘真实性’达49%;而‘真实性’的定义亦迥异——有人要求引用来源,有人主张专家意见,甚至有人期望呈现非主流观点。某些能力如拟人化行为和安全护栏,存在明显争议,部分人支持而另一些人反对。此外,用户常区分‘默认行为’与‘主动请求’场景,此类上下文差异无法被二元比较捕捉。当49%用户要求真实性却理解各异时,单个奖励模型难以实现有效对齐。即便在高投入模型中,幻觉率仍居高不下,说明现有方法未能识别真实偏好。本研究揭示了当前被简化为统一偏好模型的复杂、情境化且争议性信号,这一做法被批评为认知暴力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and values. However, this method has known limitations: it aggregates conflicting preferences, often relies on unrepresentative samples, and uses only binary comparisons. Analysing 1,500 open-ended responses from the PRISM dataset across 75 countries, we examine what people actually want from AI systems and reveal concrete failures of current methods. We find that different people want different things: most values are requested by fewer than a quarter of respondents, with truthfulness the sole exception at 49%. Furthermore, the same words hide divergent meanings: when people describe what they mean by "truthfulness", they reveal distinct, potentially incompatible, epistemological bases, as some ask for sourced claims, some for expert opinions, and some even ask for unpopular views. Certain capabilities, namely how human-like a model behaves, and some features, like AI guardrails, are outright controversial, with some desiring them and others rejecting them. We additionally find that people often use contextual distinctions (what AI should do "by default" versus "if requested") that binary comparisons cannot capture. These findings expose fundamental problems in current alignment practices. When 49% request truthfulness but define it differently, this is unlikely to be captured by a single reward model. The persistence of high hallucination rates in well-funded models, despite users' clear demands for accuracy, suggests that current methods fail to identify actual preferences. This paper sheds light on the situated, contested, imperfect signals that are currently being flattened into universal preference models, a practice others have characterised as epistemic violence.

AI对齐用户偏好多样性伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。