arXiv:2601.13669cs.CL2026-01被引 1

提出社区级对齐新范式,解决模型对少数群体价值忽视问题。

CommunityBench: Benchmarking Community-Level Alignment across Diverse Groups and Tasks

  • 基于社会群集理论设计四类任务,评估模型对群体偏好建模能力。
  • 实测主流大模型在社区偏好建模上表现有限,难以捕捉群体特异性。
  • 验证社区对齐可支撑个性化建模,为可扩展多元对齐提供新路径。

大语言模型对齐旨在使模型行为符合人类价值观。现有策略主要分为两种:一是假设存在统一的价值体系(一概而论),二是将每个个体视为独特以定制模型(个体层面)。然而,单一价值空间会边缘化少数群体规范,而为每个个体定制模型成本过高。鉴于人类社会由具有高度内部价值一致性的社会集群构成,我们提出社区级对齐作为中间方案。为此,我们构建了首个大规模社区级对齐评估基准——CommunityBench,包含基于共同身份与共同纽带理论的四项任务。通过CommunityBench,我们对多种基础模型进行了全面评估,发现当前大模型在建模社区特定偏好方面能力有限。此外,我们探讨了社区级对齐在促进个体建模中的潜力,为实现可扩展且多元化的对齐提供了新方向。

原文摘要 · Abstract (English)

Large language models (LLMs) alignment ensures model behaviors reflect human value. Existing alignment strategies primarily follow two paths: one assumes a universal value set for a unified goal (i.e., one-size-fits-all), while the other treats every individual as unique to customize models (i.e., individual-level). However, assuming a monolithic value space marginalizes minority norms, while tailoring individual models is prohibitively expensive. Recognizing that human society is organized into social clusters with high intra-group value alignment, we propose community-level alignment as a "middle ground". Practically, we introduce CommunityBench, the first large-scale benchmark for community-level alignment evaluation, featuring four tasks grounded in Common Identity and Common Bond theory. With CommunityBench, we conduct a comprehensive evaluation of various foundation models on CommunityBench, revealing that current LLMs exhibit limited capacity to model community-specific preferences. Furthermore, we investigate the potential of community-level alignment in facilitating individual modeling, providing a promising direction for scalable and pluralistic alignment.

大模型对齐社区建模价值观对齐基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。