用角色扮演量化大模型的伦理风险倾向,发现其系统性偏见
Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play
- 引入认知科学风险量表,设计新工具评估模型伦理决策风险
- 发现主流大模型在伦理决策中对不同群体存在可量化的系统偏见
- 适合关注AI安全、伦理审查与公平性的研究人员和开发者
随着大型语言模型(LLMs)日益普及,其安全性、伦理问题及潜在偏见引发关注。本研究创新性地将认知科学中的领域特定风险承担(DOSPERT)量表应用于LLMs,提出新的伦理决策风险态度量表(EDRAS),深入评估模型在伦理领域的风险态度。进一步提出融合风险量表与角色扮演的新方法,实现对LLMs系统性偏见的定量评估。通过对多个主流LLMs的系统评估,揭示了模型在多领域尤其是伦理领域的‘风险人格’特征,并量化了其对不同群体的系统性偏见。该研究有助于理解大模型的风险决策机制,保障其安全可靠应用。所提方法可作为识别与缓解偏见的工具,推动更公平可信的AI系统发展。代码与数据已公开。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) become more prevalent, concerns about their safety, ethics, and potential biases have risen. Systematically evaluating LLMs' risk decision-making tendencies and attitudes, particularly in the ethical domain, has become crucial. This study innovatively applies the Domain-Specific Risk-Taking (DOSPERT) scale from cognitive science to LLMs and proposes a novel Ethical Decision-Making Risk Attitude Scale (EDRAS) to assess LLMs' ethical risk attitudes in depth. We further propose a novel approach integrating risk scales and role-playing to quantitatively evaluate systematic biases in LLMs. Through systematic evaluation and analysis of multiple mainstream LLMs, we assessed the "risk personalities" of LLMs across multiple domains, with a particular focus on the ethical domain, and revealed and quantified LLMs' systematic biases towards different groups. This research helps understand LLMs' risk decision-making and ensure their safe and reliable application. Our approach provides a tool for identifying and mitigating biases, contributing to fairer and more trustworthy AI systems. The code and data are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。