arXiv:2603.06264cs.CLcs.CY2026-03中稿 · AAAI

检测大模型在亚洲宗教观点上的文化偏差,发现主流模型严重误判少数群体信仰。

Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinion

  • 用概率日志分析模型内部态度分布,对比真实公众意见
  • 多数模型在宗教议题上严重偏离亚洲公众观点,尤其低估少数群体
  • 轻量提示虽部分缓解偏差,但无法根治文化错配问题

大型语言模型(LLMs)在多语言、多文化场景中日益普及,但其主要基于英语数据训练,可能与不同社会的文化价值观脱节。本文对GPT-4o-Mini、Gemini-2.5-Flash、Llama 3.2、Mistral和Gemma 3在印度、东亚及东南亚的跨文化对齐性进行了全面多语言评估。研究聚焦宗教这一敏感领域,作为文化对齐的探针。通过分析模型内部表示(使用log-probs/logits),比较其观点分布与真实公众态度的差异。结果显示,尽管主流模型在一般社会议题上与公众意见基本一致,但在宗教观点上普遍失准,尤其对少数群体信仰存在系统性误判,常强化负面刻板印象。轻量级干预如人口特征提示和母语提示可部分缓解偏差,但无法消除文化鸿沟。下游偏见基准测试(如CrowS-Pairs、IndiBias、ThaiCLI、KoBBQ)进一步揭示了在敏感情境下的持续危害与代表性不足。研究强调需开展系统性、区域化扎根的审计,以保障大模型在全球部署中的公平性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly being deployed in multilingual, multicultural settings, yet their reliance on predominantly English-centric training data risks misalignment with the diverse cultural values of different societies. In this paper, we present a comprehensive, multilingual audit of the cultural alignment of contemporary LLMs including GPT-4o-Mini, Gemini-2.5-Flash, Llama 3.2, Mistral and Gemma 3 across India, East Asia and Southeast Asia. Our study specifically focuses on the sensitive domain of religion as the prism for broader alignment. To facilitate this, we conduct a multi-faceted analysis of every LLM's internal representations, using log-probs/logits, to compare the model's opinion distributions against ground-truth public attitudes. We find that while the popular models generally align with public opinion on broad social issues, they consistently fail to accurately represent religious viewpoints, especially those of minority groups, often amplifying negative stereotypes. Lightweight interventions, such as demographic priming and native language prompting, partially mitigate but do not eliminate these cultural gaps. We further show that downstream evaluations on bias benchmarks (such as CrowS-Pairs, IndiBias, ThaiCLI, KoBBQ) reveal persistent harms and under-representation in sensitive contexts. Our findings underscore the urgent need for systematic, regionally grounded audits to ensure equitable global deployment of LLMs.

文化对齐宗教偏见多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。