arXiv:2512.05176cs.SEcs.AI2025-12

构建社区文化价值观对齐评估基准,提升AI对多元文化的理解能力。

Towards A Cultural Intelligence and Values Inferences Quality Benchmark for Community Values and Common Knowledge

  • 借鉴韩国基准方法,构建面向美国社区的文化对齐评估框架
  • 聚焦社区共知与社会价值,评估LLM在多元文化中的表现差异
  • 为开发更包容的AI系统提供可量化的评测工具,适合文化敏感型AI研究者

大型语言模型(LLMs)虽被广泛应用于软件工程团队,但其设计多基于主流西方白人叙事,难以反映其他文化群体的协作创新经验。为此,学界提出如ChatBlackGPT等“文化敏感型”LLM以弥补这一偏差。然而,相关评估体系仍严重不足。现有国家对齐基准仅覆盖单一民族视角,难以反映美国多元文化现实。本文通过复现韩国国家对齐基准KorNAT的方法,构建CIVIQ——一个聚焦社区社会价值与共知知识的文化智能与价值观推断质量评估基准。该基准支持对不同社区文化背景下的模型表现进行系统评测,为实现人工智能在实践中更广泛的文化对齐提供了关键基础。

原文摘要 · Abstract (English)

Large language models (LLMs) have emerged as a powerful technology, and thus, we have seen widespread adoption and use on software engineering teams. Most often, LLMs are designed as "general purpose" technologies meant to represent the general population. Unfortunately, this often means alignment with predominantly Western Caucasian narratives and misalignment with other cultures and populations that engage in collaborative innovation. In response to this misalignment, there have been recent efforts centered on the development of "culturally-informed" LLMs, such as ChatBlackGPT, that are capable of better aligning with historically marginalized experiences and perspectives. Despite this progress, there has been little effort aimed at supporting our ability to develop and evaluate culturally-informed LLMs. A recent effort proposed an approach for developing a national alignment benchmark that emphasizes alignment with national social values and common knowledge. However, given the range of cultural identities present in the United States (U.S.), a national alignment benchmark is an ineffective goal for broader representation. To help fill this gap in this US context, we propose a replication study that translates the process used to develop KorNAT, a Korean National LLM alignment benchmark, to develop CIVIQ, a Cultural Intelligence and Values Inference Quality benchmark centered on alignment with community social values and common knowledge. Our work provides a critical foundation for research and development aimed at cultural alignment of AI technologies in practice.

文化对齐评估基准LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。