arXiv:2505.21548cs.CLcs.AI2025-05被引 10

印度本土大模型虽用印地语,却仍偏向西方文化,难符本地价值观。

Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment

  • 以印度民调和社区问答数据为基准,评估模型文化契合度。
  • 六款本土模型在价值观上不如美国用户代表的参考,且写作建议西方化。
  • 当前训练数据缺乏文化语境,需社区共建语料与深度评估体系。

大型语言模型在全球广泛应用,但普遍存在西方文化倾向。许多国家正构建‘区域性’或‘主权性’大模型,但其是否反映本地价值与实践尚不明确。本文以印度为例,评估六款印地语模型与六款全球模型在价值观与行为实践两个维度的表现,基于全国代表性调查与社区来源的问答数据集。结果显示,印地语模型在价值观上并未优于全球模型;实际上,美国受访者比任何一款印地语模型更贴近印度民众的价值观。进一步对115名印度用户开展用户研究发现,无论全球还是印地语模型生成的写作建议均呈现西方化或异域化特征。提示工程与区域微调无法恢复文化对齐,甚至可能削弱已有知识。根源在于缺乏文化语境化的预训练数据。本文将文化评估列为与多语言基准同等重要的首要要求,并提出可复用、社区驱动的评估方法,呼吁建立本地作者撰写的语料库与全面评估体系,以实现真正意义上的主权大模型。

原文摘要 · Abstract (English)

Large language models (LLMs) are used worldwide, yet exhibit Western cultural tendencies. Many countries are now building ``regional'' or ``sovereign'' LLMs, but it remains unclear whether they reflect local values and practices or merely speak local languages. Using India as a case study, we evaluate six Indic and six global LLMs on two dimensions -- values and practices -- grounded in nationally representative surveys and community-sourced QA datasets. Across tasks, Indic models do not align better with Indian norms than global models; in fact, a U.S. respondent is a closer proxy for Indian values than any Indic model. We further run a user study with 115 Indian users and find that writing suggestions from both global and Indic LLMs introduce Westernized or exoticized writing. Prompting and regional fine-tuning fail to recover alignment and can even degrade existing knowledge. We attribute this to scarce culturally grounded data, especially for pretraining. We position cultural evaluation as a first-class requirement alongside multilingual benchmarks and offer a reusable, community-grounded methodology. We call for native, community-authored corpora and thickxwide evaluations to build truly sovereign LLMs.

文化对齐大模型评估本土化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。