arXiv:2505.09595cs.CLcs.AI2025-05被引 5

测试大模型对全球文化视角的包容性,发现新方法能显著提升多元观点表达。

WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models

  • 用开放生成评估替代固定分类,衡量模型对多元文化的包容能力。
  • 多智能体协作使观点分布熵从13%升至94%,情感更积极,文化平衡改善。
  • 适合关注AI公平性、跨文化应用的研究者与开发者参考。

大型语言模型(LLMs)主要基于西方中心的知识体系和文化规范进行训练与对齐,导致文化同质化,削弱其反映全球文明多样性能力。现有评测框架因依赖僵化封闭的评估方式,难以捕捉此类偏见。为此,我们提出WorldView-Bench,一个用于评估大模型全球文化包容性(GCI)的基准,通过分析其容纳多元世界观的能力实现。该方法基于Senturk等提出的多维世界观理论,区分单一视角(Uniplex)与多维融合(Multiplex)模型。本研究采用自由生成式评估,测量文化极化程度。通过两种干预策略实现多维性:(1) 上下文嵌入多维原则的系统提示;(2) 多智能体系统(MAS)中由代表不同文化视角的多个模型协同生成。结果显示,与基线相比,多智能体实现的多维模型在观点分布得分(PDS)熵值从13%提升至94%,同时正面情感占比达67.7%,文化平衡显著增强。这些发现表明,具备多维意识的评估可有效缓解大模型中的文化偏见,推动更包容、伦理对齐的AI系统发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are predominantly trained and aligned in ways that reinforce Western-centric epistemologies and socio-cultural norms, leading to cultural homogenization and limiting their ability to reflect global civilizational plurality. Existing benchmarking frameworks fail to adequately capture this bias, as they rely on rigid, closed-form assessments that overlook the complexity of cultural inclusivity. To address this, we introduce WorldView-Bench, a benchmark designed to evaluate Global Cultural Inclusivity (GCI) in LLMs by analyzing their ability to accommodate diverse worldviews. Our approach is grounded in the Multiplex Worldview proposed by Senturk et al., which distinguishes between Uniplex models, reinforcing cultural homogenization, and Multiplex models, which integrate diverse perspectives. WorldView-Bench measures Cultural Polarization, the exclusion of alternative perspectives, through free-form generative evaluation rather than conventional categorical benchmarks. We implement applied multiplexity through two intervention strategies: (1) Contextually-Implemented Multiplex LLMs, where system prompts embed multiplexity principles, and (2) Multi-Agent System (MAS)-Implemented Multiplex LLMs, where multiple LLM agents representing distinct cultural perspectives collaboratively generate responses. Our results demonstrate a significant increase in Perspectives Distribution Score (PDS) entropy from 13% at baseline to 94% with MAS-Implemented Multiplex LLMs, alongside a shift toward positive sentiment (67.7%) and enhanced cultural balance. These findings highlight the potential of multiplex-aware AI evaluation in mitigating cultural bias in LLMs, paving the way for more inclusive and ethically aligned AI systems.

文化偏见多视角评估大模型评测多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。