评测大模型与五大国家价值观的对齐程度,发现并修复文化冲突
Benchmarking Multi-National Value Alignment for Large Language Models
- 构建多国价值观提取管道,自动筛选价值相关议题
- 在5个国家的LLM测试中发现显著价值偏差,可定位冲突场景
- 支持与对齐技术结合,提升模型在目标国家的价值一致性
大型语言模型是否违背本国价值观?有时确实如此。现有研究多聚焦伦理审查,未能涵盖政策、法律和道德等更广泛国家价值观。同时,依赖人工设计问卷的谱系测试难以扩展。为此,我们提出NaVAB,一个针对中国、美国、英国、法国和德国五个主要国家的综合性基准。NaVAB通过国家价值观提取流程,高效构建评估数据集:采用指令标注建模处理原始数据,筛选价值相关主题,并通过冲突减少机制生成非冲突值。我们在多种国家背景下的大模型上开展实验,结果揭示了误对齐场景的识别路径。此外,我们证明将NaVAB与对齐技术结合,可有效降低模型在目标国家的价值争议。
原文摘要 · Abstract (English)
Do Large Language Models (LLMs) hold positions that conflict with your country's values? Occasionally they do! However, existing works primarily focus on ethical reviews, failing to capture the diversity of national values, which encompass broader policy, legal, and moral considerations. Furthermore, current benchmarks that rely on spectrum tests using manually designed questionnaires are not easily scalable. To address these limitations, we introduce NaVAB, a comprehensive benchmark to evaluate the alignment of LLMs with the values of five major nations: China, the United States, the United Kingdom, France, and Germany. NaVAB implements a national value extraction pipeline to efficiently construct value assessment datasets. Specifically, we propose a modeling procedure with instruction tagging to process raw data sources, a screening process to filter value-related topics and a generation process with a Conflict Reduction mechanism to filter non-conflicting values.We conduct extensive experiments on various LLMs across countries, and the results provide insights into assisting in the identification of misaligned scenarios. Moreover, we demonstrate that NaVAB can be combined with alignment techniques to effectively reduce value concerns by aligning LLMs' values with the target country.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。