构建中文价值观语料库,提升大模型对主流价值的对齐能力
C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models
- 基于中国核心价值观构建分层框架,生成超25万条规则
- 70.5%测试中模型更倾向语料库生成选项,人类标注者匹配率达87.5%
- 适合关注中文伦理对齐与文化适配性的研究者使用
确保大语言模型(LLMs)与主流人类价值观和伦理规范对齐,是AI安全可持续发展的关键。现有评估与对齐方法受限于西方文化偏见及本土框架不完善,且缺乏可扩展的规则驱动场景生成手段,导致评估成本高、覆盖不足。为此,我们提出一个基于核心中国价值观的分层框架,涵盖三个主要维度、12个核心价值与50个衍生价值。基于此框架,构建了包含超过25万条价值规则的中文价值观规则语料库(C-VARC),并通过人工标注增强与扩展。实验表明,基于C-VARC生成的场景具有更清晰的价值边界和更高内容多样性。在六个敏感主题(如代孕、自杀)的评估中,七款主流大模型在70.5%以上案例中偏好C-VARC生成选项,五名中国人工标注者与语料库的对齐率达87.5%,验证其普适性、文化相关性与强对齐性。此外,我们构建了40万条基于规则的道德困境场景,客观捕捉17个大模型在价值冲突优先级上的细微差异。本工作建立了具有中国特色的文化适应性基准框架,实现全面的价值评估与对齐。
原文摘要 · Abstract (English)
Ensuring that Large Language Models (LLMs) align with mainstream human values and ethical norms is crucial for the safe and sustainable development of AI. Current value evaluation and alignment are constrained by Western cultural bias and incomplete domestic frameworks reliant on non-native rules; furthermore, the lack of scalable, rule-driven scenario generation methods makes evaluations costly and inadequate across diverse cultural contexts. To address these challenges, we propose a hierarchical value framework grounded in core Chinese values, encompassing three main dimensions, 12 core values, and 50 derived values. Based on this framework, we construct a large-scale Chinese Value Rule Corpus (C-VARC) containing over 250,000 value rules enhanced and expanded through human annotation. Experimental results demonstrate that scenarios guided by C-VARC exhibit clearer value boundaries and greater content diversity compared to those produced through direct generation. In the evaluation across six sensitive themes (e.g., surrogacy, suicide), seven mainstream LLMs preferred C-VARC generated options in over 70.5% of cases, while five Chinese human annotators showed an 87.5% alignment with C-VARC, confirming its universality, cultural relevance, and strong alignment with Chinese values. Additionally, we construct 400,000 rule-based moral dilemma scenarios that objectively capture nuanced distinctions in conflicting value prioritization across 17 LLMs. Our work establishes a culturally-adaptive benchmarking framework for comprehensive value evaluation and alignment, representing Chinese characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。