arXiv:2603.04217cs.CL2026-03Conference of the …被引 1

测试大模型对人权限制的态度,发现其存在系统性偏见。

When Do Language Models Endorse Limitations on Human Rights Principles?

  • 用1152个合成场景测试模型在24项人权条款间的权衡选择
  • 经济文化权利被允许限制的比例高于政治公民权利
  • 中文和印地语场景下更易接受权利限制,提示语言差异影响判断

随着大型语言模型(LLMs)越来越多地介入全球信息获取并可能塑造公众讨论,其与普遍人权原则的一致性变得至关重要,以确保在高风险的人工智能交互中尊重这些权利。本文通过跨8种语言、24项《世界人权宣言》(UDHR)条款的1,152个合成场景,评估了十一个主流大模型在人权权衡中的表现。分析发现,模型存在系统性偏差:(1)更倾向于接受对经济、社会和文化权利的限制,而非政治与公民权利;(2)在中文和印地语场景中,对权利限制的接受率显著高于英语或罗马尼亚语;(3)对提示工程高度敏感,易受引导;(4)李克特量表与开放回答结果存在明显差异,揭示了评估模型偏好时的关键挑战。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) increasingly mediate global information access with the potential to shape public discourse, their alignment with universal human rights principles becomes important to ensure that these rights are abided by in high stakes AI-mediated interactions. In this paper, we evaluate how LLMs navigate trade-offs involving the Universal Declaration of Human Rights (UDHR), leveraging 1,152 synthetically generated scenarios across 24 rights articles and eight languages. Our analysis of eleven major LLMs reveals systematic biases where models: (1) accept limiting Economic, Social, and Cultural rights more often than Political and Civil rights, (2) demonstrate significant cross-linguistic variation with elevated endorsement rates of rights-limiting actions in Chinese and Hindi compared to English or Romanian, (3) show substantial susceptibility to prompt-based steering, and (4) exhibit noticeable differences between Likert and open-ended responses, highlighting critical challenges in LLM preference assessment.

人权大模型偏见检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。