arXiv:2512.10955cs.CV2025-12被引 3

提出首个开放词汇的属性编码器,可精准分离图像的特定视觉属性。

Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization

  • 通过语义关联图像对与正负属性标注,训练模型学会保留或抑制特定属性
  • 在多个基准上实现属性迁移与组合生成的最先进性能
  • 适合需要精细控制图像风格、身份等属性的研究者和开发者

视觉概念个性化旨在将特定图像属性(如身份、表情、光照、风格)迁移至未见场景。然而,现有方法依赖通用图像编码器的全局嵌入,这些嵌入混杂多种视觉因素,难以分离单一属性,常导致信息泄露与合成不一致。为此,我们提出 Omni-Attribute,首个面向开放词汇的图像属性编码器,旨在学习高保真、属性特异的表示。方法上,我们联合设计数据与模型:(i) 构建语义关联的图像对,并标注正负属性,明确指导编码器应保留或抑制哪些特征;(ii) 采用双目标训练范式,在生成保真度与对比解耦之间取得平衡。所得嵌入在开放词汇属性检索、个性化及组合生成任务中表现优异,多项基准上达到当前最优性能。

原文摘要 · Abstract (English)

Visual concept personalization aims to transfer only specific image attributes, such as identity, expression, lighting, and style, into unseen contexts. However, existing methods rely on holistic embeddings from general-purpose image encoders, which entangle multiple visual factors and make it difficult to isolate a single attribute. This often leads to information leakage and incoherent synthesis. To address this limitation, we introduce Omni-Attribute, the first open-vocabulary image attribute encoder designed to learn high-fidelity, attribute-specific representations. Our approach jointly designs the data and model: (i) we curate semantically linked image pairs annotated with positive and negative attributes to explicitly teach the encoder what to preserve or suppress; and (ii) we adopt a dual-objective training paradigm that balances generative fidelity with contrastive disentanglement. The resulting embeddings prove effective for open-vocabulary attribute retrieval, personalization, and compositional generation, achieving state-of-the-art performance across multiple benchmarks.

属性编码视觉个性化解耦表示开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。