跨文化价值观影响AI安全评估,10%内容需文化敏感性标注
Quantifying the Salience of Geo-Cultural Values for Pluralistic Safety Alignment

- 用文化维度模型分析6个数据集,证明文化区归属影响评分
- 约10%的评估项具文化敏感性,需补充文化代表性
- 当前大模型不能替代人工评分,但可筛选需重点审核内容
AI模型的安全全球部署需与多元文化价值观对齐,但现有安全评估数据集的评价者多为地理同质群体,难以反映跨文化差异。进一步地,文化差异是否在控制年龄、性别、种族等人口统计因素后仍存在尚不明确。通过对多个安全评估数据集的元分析,我们发现多数数据集未报告地理文化信息,即使有也缺乏统一方法联合分析文化与人口统计因素。基于Inglehart-Welzel文化维度,多层建模显示文化区归属能解释评分方差,超越标准人口统计变量(6个数据集中均p<0.05)。分析表明,在所考察数据集中约10%的条目具有文化敏感性:若缺乏充分文化代表,可能被错误标记为安全。我们评估了大语言模型作为评分代理和筛查工具的效能,发现当前模型无法可靠替代人工评分,但可有效识别需优先进行人工标注的文化敏感项。研究呼吁更包容的文化多样性安全评估,并提供可操作建议。
原文摘要 · Abstract (English)
Safe global deployment of AI models requires alignment with human values that vary across cultures. Yet rater pools in safety evaluation datasets remain largely geographically homogeneous, failing to capture geo-cultural differences. Further, it remains unclear whether such differences persist after controlling for demographics such as age, gender, and ethnicity. Through a meta-analysis of safety datasets, we find that most do not report geo-cultural information, and those that do lack a unified methodology to jointly analyze geo-cultural and demographic correlates. Using the Inglehart-Welzel dimensions of cross-cultural variation, we demonstrate via multilevel modeling that cultural zone membership explains variance in safety ratings beyond standard demographics (p<0.05 across 6 datasets). Moreover, our analysis indicates that roughly 10% of items in the datasets we examined are culturally sensitive: likely to be misclassified as safe without adequate cultural representation. We evaluate LLMs as both rater surrogates and triage tools, finding that current LLMs do not reliably stand in for raters, though they can help prioritize culturally sensitive items for human annotation. Our findings motivate more culturally pluralistic safety evaluation and offer practical takeaways to support it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。