构建首个细粒度视觉属性数据集,实现光照、纹理等属性的精准迁移。
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models
- 将图像美学拆解为光照、纹理、动态等具体属性,支持多源属性组合
- 创建包含100万张高质量图像的FiVA数据集,标注细粒度视觉属性
- 适用于需要精细控制生成图像风格的设计师和非专业用户
文本到图像生成技术虽已取得进展,但对非专业人士而言,准确描述所需视觉属性仍具挑战。现有方法仅能从源图像中提取身份与风格,而风格涵盖范围有限,无法包含光照、动态等关键属性。为此,本文首次构建了细粒度视觉属性数据集FiVA,包含约100万张高质量生成图像及详尽的视觉属性标注。基于该数据集,提出FiVA-Adapter框架,可解耦并灵活迁移多个源图像中的具体属性至目标图像,实现更精细可控的图像生成,提升个性化定制能力。
原文摘要 · Abstract (English)
Recent advances in text-to-image generation have enabled the creation of high-quality images with diverse applications. However, accurately describing desired visual attributes can be challenging, especially for non-experts in art and photography. An intuitive solution involves adopting favorable attributes from the source images. Current methods attempt to distill identity and style from source images. However, "style" is a broad concept that includes texture, color, and artistic elements, but does not cover other important attributes such as lighting and dynamics. Additionally, a simplified "style" adaptation prevents combining multiple attributes from different sources into one generated image. In this work, we formulate a more effective approach to decompose the aesthetics of a picture into specific visual attributes, allowing users to apply characteristics such as lighting, texture, and dynamics from different images. To achieve this goal, we constructed the first fine-grained visual attributes dataset (FiVA) to the best of our knowledge. This FiVA dataset features a well-organized taxonomy for visual attributes and includes around 1 M high-quality generated images with visual attribute annotations. Leveraging this dataset, we propose a fine-grained visual attribute adaptation framework (FiVA-Adapter), which decouples and adapts visual attributes from one or more source images into a generated one. This approach enhances user-friendly customization, allowing users to selectively apply desired attributes to create images that meet their unique preferences and specific content requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。