arXiv:2510.12689cs.CYcs.AI2025-10Conference of the …被引 1

用长期利益框架优化模型,减少短视迎合,提升政策决策合理性。

From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM

  • 设计信托型模型,权衡短期与长期利益,避免盲目迎合用户偏好。
  • 在明确议题上,信托模型更贴近专家共识,准确率提升12%以上。
  • 在主观议题上,模型默认立场导致偏差,适合关注长期福祉的场景。

大型语言模型在预测问卷回答和政策偏好方面表现优异,引发对其代表人类利益潜力的关注。现有研究多聚焦于‘行为克隆’,即模型是否忠实复现个体表达的偏好。本文借鉴政治代表理论,提出关键设计权衡:AI系统应作为‘委托人’(镜像表达偏好)还是‘受托人’(判断何为个体真正利益)?该权衡与模型讨好(sycophancy)密切相关——模型可能迎合用户短期偏好,却损害其长期利益。通过一系列模拟美国政策议题投票的实验,我们采用时间效用框架,对短期与长期利益进行加权,对比信托型(侧重长期)与委托型(复制表达)模型。结果显示,信托型模型在明确议题上更接近专家共识,但在缺乏共识的议题上表现出更强的模型默认立场偏差。这揭示了代表人类利益的内在权衡:委托模型更尊重用户自主性,但可能偏离合理政策;信托模型促进长期福祉,却有家长式干预和偏见风险。

原文摘要 · Abstract (English)

Large language models (LLMs) have shown promising accuracy in predicting survey responses and policy preferences, which has increased interest in their potential to represent human interests in various domains. Most existing research has focused on "behavioral cloning", effectively evaluating how well models reproduce individuals' expressed preferences. Drawing on theories of political representation, we highlight an underexplored design trade-off: whether AI systems should act as delegates, mirroring expressed preferences, or as trustees, exercising judgment about what best serves an individual's interests. This trade-off is closely related to issues of LLM sycophancy, where models can encourage behavior or validate beliefs that may be aligned with a user's short-term preferences, but is detrimental to their long-term interests. Through a series of experiments simulating votes on various policy issues in the U.S. context, we apply a temporal utility framework that weighs short and long-term interests (simulating a trustee role) and compare voting outcomes to behavior-cloning models (simulating a delegate). We find that trustee-style predictions weighted toward long-term interests produce policy decisions that align more closely with expert consensus on well-understood issues, but also show greater bias toward models' default stances on topics lacking clear agreement. These findings reveal a fundamental trade-off in designing AI systems to represent human interests. Delegate models better preserve user autonomy but may diverge from well-supported policy positions, while trustee models can promote welfare on well-understood issues yet risk paternalism and bias on subjective topics.

AI伦理模型对齐政策推荐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。