arXiv:2605.12515cs.CL2026-05被引 1

让多语言大模型在不同语言下保持文化一致性,避免身份错乱。

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation

论文配图:Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation
图 1 · 摘自论文原文
  • 基于共识的偏好优化框架,稳定跨语言文化表现
  • 在低资源语言上提升0.13点文化一致性评分
  • 适合关注多语言模型公平性与文化安全的研究者

尽管多语言大语言模型能力强大,但当提示语种变化时,其行为常出现不一致。例如,设定英国人设后,英语提问输出莎士比亚,西班牙语提问却输出塞万提斯。为量化这种跨语言文化不一致,我们提出单一弗莱斯κ_S,该指标对幻觉具有数学鲁棒性。为此,我们提出跨语言文化一致偏好优化(C-3PO)框架。实验证明,该方法相较未对齐模型在κ_S上提升最高达0.13点,优于强提示与表征引导基线,且能保持用户身份、文化中立性与内在文化知识。评估显示,该不一致现象在印尼语、波斯语等低资源语言中尤为严重。早期解码表明,模型在前向传播中逐渐将输出个性化为提示语种的刻板文化印象。

原文摘要 · Abstract (English)

Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's language changes. While such adaptation is generally desirable, it becomes a critical failure when a user's identity is explicitly defined. For instance, given a fixed British persona and an ambiguous everyday knowledge query about literature, the prompt's language frequently overwrites the system persona -- yielding Shakespeare in English but Cervantes in Spanish. To robustly quantify this Cross-lingual Cultural Inconsistency, we introduce Singleton Fleiss's $κ_S$, a metric mathematically resilient to hallucinations. For mitigation, we propose Cross-lingual Cultural Consistent Preference Optimisation (C-3PO), a consensus-driven alignment framework. C-3PO achieves up to a 0.13-point absolute increase in $κ_S$ over unaligned models, consistently outperforming strong prompting and representation steering baselines whilst preserving explicit user identities, cultural neutrality and intrinsic cultural knowledge. Empirical evaluations demonstrate this inconsistency disproportionately affects lower-resource languages like Indonesian and Persian. Finally, early decoding of intermediate layers reveals that MLLMs implicitly personalise outputs towards the prompt language's stereotypical culture as forward-pass representations stabilise.

多语言模型文化一致性偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。