arXiv:2511.17256cs.CYcs.CL2025-11

对比中西大模型文化对齐差异,发现主流模型普遍存在价值观不稳和偏见。

Cross-cultural value alignment frameworks for responsible AI governance: Evidence from China-West comparative analysis

  • 构建四层评估平台,从伦理困境、文化忠实度等维度量化对齐效果
  • 20+模型对比显示:模型规模与对齐质量非线性相关,且年轻群体代表性不足
  • 中国模型重多语言优化,西方模型有架构创新但仍存美国中心偏见

随着大型语言模型(LLMs)在跨文化高风险决策中的作用日益增强,确保其与多元文化价值对齐已成为关键治理挑战。本研究提出一个多层负责任AI审计平台,通过四种集成方法系统评估中国源与西方源大模型的跨文化价值对齐:用于评估时间稳定性的伦理困境语料库、用于量化文化忠实度的多样性增强框架(DEF)、用于分布准确性的首词概率对齐,以及用于可解释决策的多阶段推理框架(MARK)。对20多个领先模型(如Qwen、GPT-4o、Claude、LLaMA、DeepSeek)的比较分析揭示了普遍性问题:价值观体系基础不稳、年轻群体系统性代表性不足,以及模型规模与对齐质量之间存在非线性关系;同时呈现出显著区域发展轨迹差异。中国源模型更强调多语言数据融合以实现情境优化,而西方模型虽具更大架构实验性,却持续存在美国中心偏见。两种范式均未实现稳健的跨文化泛化能力。研究发现,Mistral系列架构在跨文化对齐上显著优于LLaMA3系列,且全参数微调在多样化数据集上的表现优于基于人类反馈的强化学习。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) increasingly influence high-stakes decision-making across global contexts, ensuring their alignment with diverse cultural values has become a critical governance challenge. This study presents a Multi-Layered Auditing Platform for Responsible AI that systematically evaluates cross-cultural value alignment in China-origin and Western-origin LLMs through four integrated methodologies: Ethical Dilemma Corpus for assessing temporal stability, Diversity-Enhanced Framework (DEF) for quantifying cultural fidelity, First-Token Probability Alignment for distributional accuracy, and Multi-stAge Reasoning frameworK (MARK) for interpretable decision-making. Our comparative analysis of 20+ leading models, such as Qwen, GPT-4o, Claude, LLaMA, and DeepSeek, reveals universal challenges-fundamental instability in value systems, systematic under-representation of younger demographics, and non-linear relationships between model scale and alignment quality-alongside divergent regional development trajectories. While China-origin models increasingly emphasize multilingual data integration for context-specific optimization, Western models demonstrate greater architectural experimentation but persistent U.S.-centric biases. Neither paradigm achieves robust cross-cultural generalization. We establish that Mistral-series architectures significantly outperform LLaMA3-series in cross-cultural alignment, and that Full-Parameter Fine-Tuning on diverse datasets surpasses Reinforcement Learning from Human Feedback in preserving cultural variation...

AI治理文化对齐大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。