提出身份协商机制,让生成式AI在文化判断中更公正地处理身份问题。
From Bias Mitigation to Bias Negotiation: Governing Identity and Sociocultural Reasoning in Generative AI
- 将身份治理从消除偏见转向动态协商,强调情境化文化推理
- 发现模型常用概率表述群体倾向与权衡伤害价值来应对争议
- 适合关注AI伦理、社会公平与跨文化交互的研究者与开发者
大型语言模型通过依赖共享文化模式来理解并回应社会情境。由于身份常构成合理判断的推理基础,伦理对齐要求规范系统何时及如何调用身份信息。当前主流的身份治理仍以偏见缓解为主,将身份视为可量化的差异或有害关联,需检测并压制。这忽视了身份在解释中的积极、情境敏感作用。本文提出‘偏见协商’概念:对身份相关社会文化相关性、推断与正当性判断进行规范调控。我们通过半结构化访谈多个公开部署聊天机器人,探索其可行性,识别出重复出现的协商策略,包括对群体倾向的概率化表述和伤害价值平衡。同时观察到模型在回避关键权衡或原则应用不一致时的失效模式。偏见协商关乎正义,因正向的文化推理能力是识别并可能修复结构性不公所必需。它也直接影响模型核心功能,因跨异质文化环境运行需要社会文化能力。由于偏见协商是一种通过对话与审议表达的程序性能力,无法仅靠静态基准验证。为此,我们提出一个明确框架,将偏见协商分解为可观察与评分的协商动作空间,以及协商对象特征集,支持系统性测试套件设计与评估。
原文摘要 · Abstract (English)
LLMs act in the social world by drawing upon shared cultural patterns to make social situations understandable and actionable. Because identity is often part of the inferential substrate of competent judgment, ethical alignment requires regulating when and how systems invoke identity. Yet the dominant governance regime for identity-related harm remains bias mitigation, which treats identity primarily as a source of measurable disparities or harmful associations to be detected and suppressed. This leaves underspecified a positive, context-sensitive role for identity in interpretation. We call this governance problem bias negotiation: the normative regulation of identity-conditioned judgments of sociocultural relevance, inference, and justification. Empirically, we probe the feasibility of bias negotiation through semi-structured interviews with multiple publicly deployed chatbots. We identify recurring repertoires for negotiating identity including probabilistic framing of group tendencies and harm-value balancing. We also observe failure modes in which models avoid hard tradeoffs or apply principles inconsistently. Bias negotiation matters for justice because a positive role for sociocultural reasoning is required to recognize and potentially remediate structural inequities. But it is equally implicated in core model functionality as sociocultural competence is needed for systems that operate across heterogeneous cultural contexts. Because bias negotiation is a procedural capability expressed through deliberation and interaction, it cannot be validated by static benchmarks alone. To support targeted training, we introduce a broad but explicit framework that decomposes bias negotiation into an action space of negotiation moves (what to observe and score) and a complementary set of case features (over which the model negotiates), enabling systematic test-suite design and evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。