arXiv:2410.17032cs.AI2024-10被引 6

研究不同人群对AI生成图像安全性的判断差异,发现认知大不相同。

Insights on Disagreement Patterns in Multimodal Safety Perception across Diverse Rater Groups

  • 630名跨年龄性别族裔的参与者并行评估1000张AI图像的安全性
  • 不同群体对危害严重性的判断差异显著,且与具体违规类型相关
  • 普通人判断模式不同于专家,且与文本安全评估结果不同

生成式AI的安全评估高度依赖人类标注,但现有方法常将评分聚合,掩盖了真实世界中多元视角的差异。尤其在多模态安全领域,不同社会文化背景下的感知和伤害认知存在显著差异。本研究通过大规模实验,招募630名来自30个交叉身份组(年龄、性别、族裔)的参与者,对约1000张文本到图像生成内容进行并行安全评级。结果显示:(1)不同群体对危害严重性的评估存在显著差异,且差异随安全违规类型变化;(2)普通人群的标注模式与受特定安全政策训练的专家明显不同;(3)这些图像安全判断差异不同于以往发现的文本安全任务中的群体差异。进一步的开放性解释分析揭示了各群体对危害感知的根本原因差异。研究强调,生成式AI的安全评估必须纳入多样化视角,才能真正实现包容性与用户价值契合。

原文摘要 · Abstract (English)

AI systems crucially rely on human ratings, but these ratings are often aggregated, obscuring the inherent diversity of perspectives in real-world phenomenon. This is particularly concerning when evaluating the safety of generative AI, where perceptions and associated harms can vary significantly across socio-cultural contexts. While recent research has studied the impact of demographic differences on annotating text, there is limited understanding of how these subjective variations affect multimodal safety in generative AI. To address this, we conduct a large-scale study employing highly-parallel safety ratings of about 1000 text-to-image (T2I) generations from a demographically diverse rater pool of 630 raters balanced across 30 intersectional groups across age, gender, and ethnicity. Our study shows that (1) there are significant differences across demographic groups (including intersectional groups) on how severe they assess the harm to be, and that these differences vary across different types of safety violations, (2) the diverse rater pool captures annotation patterns that are substantially different from expert raters trained on specific set of safety policies, and (3) the differences we observe in T2I safety are distinct from previously documented group level differences in text-based safety tasks. To further understand these varying perspectives, we conduct a qualitative analysis of the open-ended explanations provided by raters. This analysis reveals core differences into the reasons why different groups perceive harms in T2I generations. Our findings underscore the critical need for incorporating diverse perspectives into safety evaluation of generative AI ensuring these systems are truly inclusive and reflect the values of all users.

多模态安全人因评估生成式AI多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。