arXiv:2506.17367cs.CLcs.AI2025-06被引 1

测试大模型对用户不便的金钱估值,发现结果混乱且不合理。

Cash or Comfort? How LLMs Value Your Inconvenience

  • 让多个大模型为走路、等待等不适定价,观察其决策逻辑。
  • 同一模型对相似问题回答差异大,有时1欧元换10小时等待。
  • 模型会拒绝无代价的高额报酬,或接受极低补偿忍受严重不适。

大型语言模型(LLMs)正被视作能代表人类做出日常决策的近自主人工智能代理。尽管在诸多技术任务中表现良好,其在个人决策行为仍不明确。现有研究多关注理性与道德对齐,但未深入探讨当金钱收益与用户舒适度冲突时的表现。本文通过量化多个大模型对一系列用户不适(额外步行、等待、饥饿、疼痛)所赋予的价格,揭示了若干关键问题:(1)不同模型间响应差异显著;(2)同一模型对提示微小变化敏感(如改用第一人称提问可大幅改变决策);(3)模型可接受极低回报应对重大不便(如1欧元换取10小时等待);(4)即使无任何不适,模型也会拒绝高额报酬(如1000欧元换取0分钟等待)。这些发现凸显出必须审慎评估当前大模型对人类不便的价值判断,尤其是在其代为决策的应用场景中。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly proposed as near-autonomous artificial intelligence (AI) agents capable of making everyday decisions on behalf of humans. Although LLMs perform well on many technical tasks, their behaviour in personal decision-making remains less understood. Previous studies have assessed their rationality and moral alignment with human decisions. However, the behaviour of AI assistants in scenarios where financial rewards are at odds with user comfort has not yet been thoroughly explored. In this paper, we tackle this problem by quantifying the prices assigned by multiple LLMs to a series of user discomforts: additional walking, waiting, hunger and pain. We uncover several key concerns that strongly question the prospect of using current LLMs as decision-making assistants: (1) a large variance in responses between LLMs, (2) within a single LLM, responses show fragility to minor variations in prompt phrasing (e.g., reformulating the question in the first person can considerably alter the decision), (3) LLMs can accept unreasonably low rewards for major inconveniences (e.g., 1 Euro to wait 10 hours), and (4) LLMs can reject monetary gains where no discomfort is imposed (e.g., 1,000 Euro to wait 0 minutes). These findings emphasize the need for scrutiny of how LLMs value human inconvenience, particularly as we move toward applications where such cash-versus-comfort trade-offs are made on users' behalf.

大模型决策人性价值行为测试智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。