区分语言模型中的文化信号与表面线索,避免误判偏见。
Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

- 设计多智能体审计框架,分离偏见、身份表示与跨文化模式三类问题。
- 9万条输出显示,英语和中文中移除身份线索后预测准确率大幅下降,法语影响小。
- 翻译和遮蔽姓名后,源语言识别率显著降低,说明表面线索易被误当文化理解。
多语言大模型的输出会因社会文化背景而异。但文化关联性证据可能具有误导性:身份标签可通过显性或隐性文本线索推断,姓名和用词也能暴露源语言。若将所有这些信号视为文化根基,可能掩盖潜在偏见。本文提出一种经人工验证的多智能体审计方法,区分三个问题:输出是否再现社会偏见、不同身份群体是否被差异化呈现、输出是否反映跨文化规律。研究分析了12个模型在英语、法语和中文下共89,253条输出,涵盖18种职业及三种任务条件。结果发现,偏见表现随语言和任务系统性变化。在英语和中文中,去除直接身份线索后,身份标签预测准确率显著下降,但在法语中影响较小。所有语言-体裁组合中,源语言的文化背景平均相关性最高,自动评分与人工评分具中等一致性。然而,翻译后及遮蔽姓名后,源语言识别率大幅下降。若不控制这些因素,多语言审计可能将表面线索误认为文化理解,导致对跨文化差异与偏见的错误结论。本审计提供了一套实用框架,帮助区分此类捷径与真正有意义的跨文化模式。
原文摘要 · Abstract (English)
Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language. Treating all these signals as evidence of cultural grounding may obscure potential biases. We present a human-validated, multi-agent audit that separates three questions: whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect cross-cultural patterns. The study analyzes 89,253 outputs from 12 LLMs in English, French, and Chinese, spanning 18 occupations and three task conditions. We find that bias representation varies systematically across languages and tasks. Removing direct identity cues sharply reduces identity-label prediction in English and Chinese, but has a much smaller effect in French. Across all language-genre settings, the cultural context associated with the source language receives the highest average relevance score, with moderate agreement between automated and human ratings. However, the ability to identify the source language drops substantially after translation and again after masking names. Without these controls, multilingual audits may mistake surface cues for cultural understanding, leading to misleading conclusions about cross-cultural variation and bias. Our audit offers a practical framework for separating such shortcuts from more meaningful cross-cultural patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。