arXiv:2412.06483cs.CLcs.AI2024-12NeurIPS被引 13

构建跨文化法律安全基准,提升大模型全球合规能力

SafeWorld: Geo-Diverse Safety Alignment

  • 设计包含50国493区域的文化与法律数据集,评估模型跨文化响应能力
  • 自建多维评估框架,发现主流模型在恰当性、准确性上均不达标
  • 通过偏好对训练,使模型在人类评测中帮助性与安全性提升近20%

在大语言模型快速发展的背景下,确保安全至关重要,但现有研究常忽视全球文化与法律标准的地理差异。为揭示这一挑战,我们提出SafeWorld——一个专门用于评估模型在多样全球语境下生成既有用又符合文化敏感性与法律合规性响应的新基准。该基准包含2,342个经人工验证的高质量测试查询,覆盖50个国家及493个地区/族裔的文化规范与法律政策。在此基础上,我们构建了一个多维度自动安全评估框架,用于衡量响应的上下文适宜性、准确性和完整性。评估结果显示,当前主流模型在这些维度上表现不佳。为此,我们采用直接偏好优化(DPO)对齐训练,合成有助于提升行为恰当性的偏好对,并鼓励模型在必要时提供具体的文化规范与政策依据。经训练的SafeWorldLM在所有三项评估维度上显著优于包括GPT-4o在内的竞品模型,全球人工评估亦显示其在帮助性与有害性判断中胜率高出近20%。代码与数据已开源:https://github.com/PlusLabNLP/SafeWorld。

原文摘要 · Abstract (English)

In the rapidly evolving field of Large Language Models (LLMs), ensuring safety is a crucial and widely discussed topic. However, existing works often overlook the geo-diversity of cultural and legal standards across the world. To demonstrate the challenges posed by geo-diverse safety standards, we introduce SafeWorld, a novel benchmark specifically designed to evaluate LLMs' ability to generate responses that are not only helpful but also culturally sensitive and legally compliant across diverse global contexts. SafeWorld encompasses 2,342 test user queries, each grounded in high-quality, human-verified cultural norms and legal policies from 50 countries and 493 regions/races. On top of it, we propose a multi-dimensional automatic safety evaluation framework that assesses the contextual appropriateness, accuracy, and comprehensiveness of responses. Our evaluations reveal that current LLMs struggle to meet these criteria. To enhance LLMs' alignment with geo-diverse safety standards, we synthesize helpful preference pairs for Direct Preference Optimization (DPO) alignment training. The preference pair construction aims to encourage LLMs to behave appropriately and provide precise references to relevant cultural norms and policies when necessary. Our trained SafeWorldLM outperforms all competing models, including GPT-4o on all three evaluation dimensions by a large margin. Global human evaluators also note a nearly 20% higher winning rate in helpfulness and harmfulness evaluation. Our code and data can be found here: https://github.com/PlusLabNLP/SafeWorld.

大模型安全跨文化法律合规评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。