arXiv:2507.13383cs.LGcs.AI2025-07NeurIPS被引 13

构建首个多元安全对齐数据集,让文生图模型理解不同人群的多样性安全观。

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

  • 设计跨人口统计学的多视角评估数据集,覆盖1000个提示词的深度反馈。
  • 发现不同群体对危害的感知差异显著,传统评估方法难以捕捉这种复杂性。
  • 适合关注公平性、可解释性和价值观对齐的研究者与开发者。

当前文生图(T2I)模型常忽视多样化的个人经验,导致系统对齐失败。本文倡导多元对齐,即让AI理解并可引导至多种相互冲突的人类价值观。研究贡献包括:第一,提出首个用于多元对齐的多模态数据集——多样交叉视觉评估数据集(DIVE),通过大量具有交叉人口统计特征的人类评估者对1000个提示词提供详尽反馈,实现高复现性,捕捉细微的安全感知差异;第二,实证表明人口统计学是该领域多样观点的关键代理变量,揭示了显著且情境依赖的危害感知差异,与传统评估结果明显不同;第三,讨论构建对齐T2I模型的启示,包括高效数据收集策略、大语言模型判断能力及模型对多元视角的可调控性。本研究为更公平、更对齐的T2I系统提供基础工具。内容警告:论文包含可能有害的敏感内容。

原文摘要 · Abstract (English)

Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enable deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems. Content Warning: The paper includes sensitive content that may be harmful.

文生图价值观对齐公平性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。