arXiv:2507.16033cs.HCcs.AI2025-07中稿 · AAAI被引 1

AI图像安全评估中,人工标注者依赖道德与社会经验判断,远超标准分类。

"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives

  • 分析5372条开放评论,发现标注者基于道德、情感与情境综合判断安全
  • 标注指南直接影响判断结果,连带改变对危害的道德解释
  • 现有框架忽略画质偏差与提示不符等主观感知因素,需改进评估设计

理解生成式AI内容的安全性定义复杂。开发者常依赖预设分类体系,但现实中的安全判断还涉及个人、社会与文化层面的伤害感知。本文通过分析5372条开放评论,揭示标注者在评估AI生成图像安全性时,持续运用道德、情感与情境推理,超越结构化安全类别。多数人关注对他人可能造成的伤害,而非自身,其判断根植于生活经验、集体风险意识与社会文化认知。此外,任务结构本身——特别是标注指南——深刻影响标注者对伤害的理解与表达。指南不仅决定哪些图像被标记,也塑造了背后道德理由。标注者频繁提及图像质量、视觉畸变及提示与输出不一致等因素作为感知危害的依据,这些在现有评估框架中常被忽视。研究显示,当前安全流程遗漏了标注者带来的关键推理维度。我们主张采用能引导道德反思、区分伤害类型、容纳主观与情境敏感解读的评估设计。

原文摘要 · Abstract (English)

Understanding what constitutes safety in AI-generated content is complex. While developers often rely on predefined taxonomies, real-world safety judgments also involve personal, social, and cultural perceptions of harm. This paper examines how annotators evaluate the safety of AI-generated images, focusing on the qualitative reasoning behind their judgments. Analyzing 5,372 open-ended comments, we find that annotators consistently invoke moral, emotional, and contextual reasoning that extends beyond structured safety categories. Many reflect on potential harm to others more than to themselves, grounding their judgments in lived experience, collective risk, and sociocultural awareness. Beyond individual perceptions, we also find that the structure of the task itself -- including annotation guidelines -- shapes how annotators interpret and express harm. Guidelines influence not only which images are flagged, but also the moral judgment behind the justifications. Annotators frequently cite factors such as image quality, visual distortion, and mismatches between prompt and output as contributing to perceived harm dimensions, which are often overlooked in standard evaluation frameworks. Our findings reveal that existing safety pipelines miss critical forms of reasoning that annotators bring to the task. We argue for evaluation designs that scaffold moral reflection, differentiate types of harm, and make space for subjective, context-sensitive interpretations of AI-generated content.

AI安全人类评估道德判断图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。