arXiv:2607.21757cs.CYcs.AI2026-07

让用户参与设计AI偏好代理,反而可能掩盖其与真实偏好的偏差。

Co-design of LLM-based preference agents: participation may drive overtrust

  • 12人共同设计家庭能源偏好代理,通过调研与访谈提升参与感
  • 人工验证显示代理响应更同质、决断且抽象,与人类样本对齐度有限
  • 参与过程透明反成过度信任诱因,需警惕规模化应用的结构性风险

大语言模型被越来越多地用于模拟人类偏好,引发验证性、误表征与排除性问题。与所代表人群共同设计代理是缓解这些问题的潜在途径,但参与本身也可能掩盖其表面解决的问题。本文通过一项以12名参与者为主的定性研究,在家庭能源领域开展背景调查、共设计访谈与验证调查,探讨该张力。参与者积极参与,普遍认为代理能良好代表自己。然而独立验证显示,代理响应在整体上更为同质、决断且抽象,与人类样本的对齐程度不一。本文提出,参与和过程透明可能成为‘过度信任引擎’,在增强信任的同时隐藏系统性偏差,具有规模化应用的潜在结构风险。为此,将个体对齐视为动态建构过程而非静态状态,构建参与式偏好代理设计的核心机制。

原文摘要 · Abstract (English)

Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an "overtrust engine" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.

LLM偏好建模共设计信任机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。