用少量例子构建可操作的视觉情绪空间,让抽象概念变直观。
"I Know It When I See It": Mood Spaces for Connecting and Expressing Visual Concepts
- 通过示例构建情绪空间,自动捕捉图像间的语义关联。
- 仅需2-20个样本,1分钟内完成学习,空间压缩50-100倍。
- 适合创意设计、跨模态生成等需要表达抽象概念的场景。
表达可标注的概念很简单,但许多想法难以定义却一眼即懂。我们提出一种情绪板(Mood Board),用户通过示例暗示属性变化的方向。计算出一个底层的情绪空间,能分离无关特征并建立图像间联系,使相关概念在空间中更接近。我们引入纤维化计算,将预训练特征压缩/解压至50-100倍更小的紧凑空间。核心创新在于学习模仿示例间图像标记的成对亲和关系。为聚焦情绪空间中的粗粒度到细粒度层次结构,我们从亲和矩阵中计算主特征向量,并在该空间定义损失函数。最终的情绪空间局部线性且紧凑,支持图像级操作,如对象平均、视觉类比和姿态迁移,均可在情绪空间中通过简单向量运算实现。学习过程计算高效,无需微调,仅需少量(2-20)示例,且学习时间不足一分钟。
原文摘要 · Abstract (English)
Expressing complex concepts is easy when they can be labeled or quantified, but many ideas are hard to define yet instantly recognizable. We propose a Mood Board, where users convey abstract concepts with examples that hint at the intended direction of attribute changes. We compute an underlying Mood Space that 1) factors out irrelevant features and 2) finds the connections between images, thus bringing relevant concepts closer. We invent a fibration computation to compress/decompress pre-trained features into/from a compact space, 50-100x smaller. The main innovation is learning to mimic the pairwise affinity relationship of the image tokens across exemplars. To focus on the coarse-to-fine hierarchical structures in the Mood Space, we compute the top eigenvector structure from the affinity matrix and define a loss in the eigenvector space. The resulting Mood Space is locally linear and compact, allowing image-level operations, such as object averaging, visual analogy, and pose transfer, to be performed as a simple vector operation in Mood Space. Our learning is efficient in computation without any fine-tuning, needs only a few (2-20) exemplars, and takes less than a minute to learn.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。