构建高质量纹理数据集,解决机器学习中纹理偏见研究的数据瓶颈。
On Synthetic Texture Datasets: Challenges, Creation, and Curation
- 用文本提示+Stable Diffusion生成纹理图像,再经多轮筛选提升质量。
- 创建含24.6万张图像的PTD数据集,覆盖56类纹理,质量获人工验证。
- 发现生成模型安全过滤器对纹理敏感(最高60%误判),揭示潜在偏差。
纹理对机器学习模型的影响长期受到关注,尤其在纹理偏差、可解释性和鲁棒性方面。然而,受限于大规模多样化的纹理数据,现有研究难以开展全面评估。图像生成模型虽可规模化生成数据,但用于纹理合成仍处于空白,且面临生成准确性与图像验证双重挑战。本文提出一种可扩展的方法,构建高质量、多样化的纹理图像数据集以支持各类纹理相关任务。该流程包括:(1) 从多种描述符生成文本提示输入至文生图模型;(2) 采用并优化Stable Diffusion管道生成并筛选图像;(3) 进一步精炼至最高质量图像。最终生成Prompted Textures Dataset(PTD),包含246,285张图像,覆盖56类纹理。在生成过程中发现,图像生成管道中的NSFW安全过滤器对纹理极为敏感,高达60%的纹理图像被误标,暴露出模型潜在偏差,为纹理数据处理带来独特挑战。通过标准指标与人工评估,确认本数据集具有高质与多样性。数据集已公开于https://zenodo.org/records/15359142。
原文摘要 · Abstract (English)
The influence of textures on machine learning models has been an ongoing investigation, specifically in texture bias/learning, interpretability, and robustness. However, due to the lack of large and diverse texture data available, the findings in these works have been limited, as more comprehensive evaluations have not been feasible. Image generative models are able to provide data creation at scale, but utilizing these models for texture synthesis has been unexplored and poses additional challenges both in creating accurate texture images and validating those images. In this work, we introduce an extensible methodology and corresponding new dataset for generating high-quality, diverse texture images capable of supporting a broad set of texture-based tasks. Our pipeline consists of: (1) developing prompts from a range of descriptors to serve as input to text-to-image models, (2) adopting and adapting Stable Diffusion pipelines to generate and filter the corresponding images, and (3) further filtering down to the highest quality images. Through this, we create the Prompted Textures Dataset (PTD), a dataset of 246,285 texture images that span 56 textures. During the process of generating images, we find that NSFW safety filters in image generation pipelines are highly sensitive to texture (and flag up to 60\% of our texture images), uncovering a potential bias in these models and presenting unique challenges when working with texture data. Through both standard metrics and a human evaluation, we find that our dataset is high quality and diverse. Our dataset is available for download at https://zenodo.org/records/15359142.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。