arXiv:2411.08243cs.CLcs.CY2024-11NAACL被引 10

审计主流安全数据集,发现其缺陷可能加剧模型对不同人群的不公平风险。

Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset

  • 通过人工与自动化方式全面评估HH数据集内容质量
  • 揭示该数据集导致模型在不同群体间产生差异化的安全行为
  • 适合关注大模型安全评估与公平性的研究者阅读

为缓解大语言模型(LLMs)带来的危害,学习人类反馈(LHF)被广泛用于引导模型输出更无害且更有帮助的内容。尽管LHF在实践中普遍采用,但其反馈质量及作为安全缓解手段的有效性仍不明确。本研究针对Anthropic发布的广泛使用的Helpful and Harmless(HH)数据集进行审计,包括:(1) 通过人工与自动评估深入分析数据集内容;(2) 实验验证该数据集对模型安全性的影响;(3) 分析引用该数据集的100篇最具影响力论文。审计结果显示,HH数据集中存在的概念化失误与质量缺陷,可能导致模型在不同人口群体间表现出不同的安全行为,从而引发新的危害。研究强调需要更精细、情境敏感的安全缓解策略。

原文摘要 · Abstract (English)

In an effort to mitigate the harms of large language models (LLMs), learning from human feedback (LHF) has been used to steer LLMs towards outputs that are intended to be both less harmful and more helpful. Despite the widespread adoption of LHF in practice, the quality of this feedback and its effectiveness as a safety mitigation technique remain unclear. This study addresses these issues by auditing the widely-used Helpful and Harmless (HH) dataset by Anthropic. Our work includes: (1) a thorough investigation of the dataset's content through both manual and automated evaluation; (2) experiments demonstrating the dataset's impact on models' safety; and (3) an analysis of the 100 most influential papers citing this dataset. Through our audit, we showcase how conceptualization failures and quality issues identified in the HH dataset can create additional harms by leading to disparate safety behaviors across demographic groups. Our findings highlight the need for more nuanced, context-sensitive approaches to safety mitigation in LLMs.

大模型安全数据审计公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。