arXiv:2505.00751cs.CV2025-05被引 2

用自然语言精准修改图像中物体的颜色和材质,同时保持结构不变。

InstructAttribute: Fine-grained Object Attributes editing with Instruction

  • 通过调整扩散模型的注意力机制实现细粒度属性编辑
  • 在多种物体类别上准确修改颜色与材质,保持图像整体一致性
  • 适合产品设计、电商展示等需要高精度图像编辑的场景

文本到图像(T2I)扩散模型因其强大的生成能力被广泛用于图像编辑,但对特定物体属性(如颜色、材质)实现精细控制仍具挑战。现有方法常难以准确修改属性或损害图像结构与整体一致性。为此,我们提出无需训练的SPAA框架,通过智能操控扩散模型中的自注意力图和交叉注意力值,实现同一物体颜色与材质的精确生成。基于SPAA,我们融合多模态大语言模型自动化构建数据集并生成指令,形成涵盖多种物体类别下丰富颜色与材质的属性数据集。利用该数据集,我们提出InstructAttribute——一个指令微调模型,可通过自然语言提示实现对象级别的细粒度属性编辑。大量实验表明,InstructAttribute优于现有基于指令的基线方法,在属性修改准确性与结构保留之间取得更优平衡。

原文摘要 · Abstract (English)

Text-to-image (T2I) diffusion models are widely used in image editing due to their powerful generative capabilities. However, achieving fine-grained control over specific object attributes, such as color and material, remains a considerable challenge. Existing methods often fail to accurately modify these attributes or compromise structural integrity and overall image consistency. To fill this gap, we introduce Structure Preservation and Attribute Amplification (SPAA), a novel training-free framework that enables precise generation of color and material attributes for the same object by intelligently manipulating self-attention maps and cross-attention values within diffusion models. Building on SPAA, we integrate multi-modal large language models (MLLMs) to automate data curation and instruction generation. Leveraging this object attribute data collection engine, we construct the Attribute Dataset, encompassing a comprehensive range of colors and materials across diverse object categories. Using this generated dataset, we propose InstructAttribute, an instruction-tuned model that enables fine-grained and object-level attribute editing through natural language prompts. This capability holds significant practical implications for diverse fields, from accelerating product design and e-commerce visualization to enhancing virtual try-on experiences. Extensive experiments demonstrate that InstructAttribute outperforms existing instruction-based baselines, achieving a superior balance between attribute modification accuracy and structural preservation.

图像编辑属性控制扩散模型自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。