研究大模型如何模仿人类风险偏好,发现其存在偏差需优化设计。
Can Risk-taking AI-Assistants suitably represent entities
- 通过多场景测试大模型的风险行为,评估其对人类偏好的复制能力。
- 发现DeepSeek Reasoner和Gemini-2.0-flash-lite在部分场景中与人类一致,但仍有显著差异。
- 适合关注AI伦理、风险决策与模型可解释性的研究人员参考。
负责任的AI要求系统的行为倾向可测量、可审计并可调整,以防止无意中引导用户做出高风险决策或隐藏风险规避偏见。随着语言模型(LMs)越来越多地融入人工智能决策支持系统,理解其风险行为对负责任部署至关重要。本研究探讨了语言模型中风险规避的可操控性(MoRA),考察其在多种经济情境下复制人类风险偏好的能力,重点关注性别差异、不确定性、角色化决策及风险规避的可操控性。结果表明,尽管DeepSeek Reasoner和Gemini-2.0-flash-lite在某些情况下表现出与人类行为的部分一致性,但显著差异凸显了需要改进以生物为中心的可操控性度量方法。研究建议应进一步优化模型设计,使人工智能更准确地复现人类风险偏好,从而提升其在风险管理中的有效性。该方法有助于增强AI助手在风险管控中的适用性。
原文摘要 · Abstract (English)
Responsible AI demands systems whose behavioral tendencies can be effectively measured, audited, and adjusted to prevent inadvertently nudging users toward risky decisions or embedding hidden biases in risk aversion. As language models (LMs) are increasingly incorporated into AI-driven decision support systems, understanding their risk behaviors is crucial for their responsible deployment. This study investigates the manipulability of risk aversion (MoRA) in LMs, examining their ability to replicate human risk preferences across diverse economic scenarios, with a focus on gender-specific attitudes, uncertainty, role-based decision-making, and the manipulability of risk aversion. The results indicate that while LMs such as DeepSeek Reasoner and Gemini-2.0-flash-lite exhibit some alignment with human behaviors, notable discrepancies highlight the need to refine bio-centric measures of manipulability. These findings suggest directions for refining AI design to better align human and AI risk preferences and enhance ethical decision-making. The study calls for further advancements in model design to ensure that AI systems more accurately replicate human risk preferences, thereby improving their effectiveness in risk management contexts. This approach could enhance the applicability of AI assistants in managing risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。