arXiv:2503.01894cs.CVcs.AI2025-03ICML被引 15

构建包容性城市空间的多标准对齐数据集,让AI图像生成更贴近社区真实需求。

LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces

  • 通过30个社区组织两年协作,构建含3.7万组对比的视觉空间数据集。
  • 用偏好优化微调模型后,高标注量下与人工偏好对齐度显著提升。
  • 揭示不同群体在安全、包容等维度有系统性差异,适合城市规划研究者使用。

我们提出本地交叉视角视觉空间(LIVS)数据集,作为多标准对齐的基准,历时两年与30个社区组织合作开发,旨在支持文本到图像(T2I)模型在包容性城市规划中的多元对齐。该数据集包含13,462张图像的37,710组成对比较,基于634个社区定义的概念,涵盖可及性、安全性、舒适度、吸引力、包容性和多样性六项标准。采用直接偏好优化(DPO)微调Stable Diffusion XL模型,通过四个案例研究验证:(1)当标注量高时,DPO显著提升与标注偏好的对齐;(2)不同身份参与者偏好模式各异,凸显交叉性数据必要性;(3)人工撰写提示生成更具区分性的视觉输出,影响标注决断力;(4)交叉群体在各标准上系统性评分差异,暴露单一目标对齐的局限。尽管DPO在特定条件下改善对齐,但中性评分普遍,表明社区价值多样且常模糊。LIVS为开发融入本地、利益相关者驱动偏好的T2I模型提供基准,奠定空间设计中情境感知对齐的基础。

原文摘要 · Abstract (English)

We introduce the Local Intersectional Visual Spaces (LIVS) dataset, a benchmark for multi-criteria alignment, developed through a two-year participatory process with 30 community organizations to support the pluralistic alignment of text-to-image (T2I) models in inclusive urban planning. The dataset encodes 37,710 pairwise comparisons across 13,462 images, structured along six criteria - Accessibility, Safety, Comfort, Invitingness, Inclusivity, and Diversity - derived from 634 community-defined concepts. Using Direct Preference Optimization (DPO), we fine-tune Stable Diffusion XL to reflect multi-criteria spatial preferences and evaluate the LIVS dataset and the fine-tuned model through four case studies: (1) DPO increases alignment with annotated preferences, particularly when annotation volume is high; (2) preference patterns vary across participant identities, underscoring the need for intersectional data; (3) human-authored prompts generate more distinctive visual outputs than LLM-generated ones, influencing annotation decisiveness; and (4) intersectional groups assign systematically different ratings across criteria, revealing the limitations of single-objective alignment. While DPO improves alignment under specific conditions, the prevalence of neutral ratings indicates that community values are heterogeneous and often ambiguous. LIVS provides a benchmark for developing T2I models that incorporate local, stakeholder-driven preferences, offering a foundation for context-aware alignment in spatial design.

文本生成图像城市规划数据集包容性设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。