让机器自动发现真实图像中的语义概念轴并精准控制。
Bridging the gap to real-world language-grounded visual concept learning
- 用通用提示策略自动发现图像语义轴,无需预设
- 在真实数据集上实现多样概念的独立编辑
- 适合需要可扩展视觉概念控制的研究者
人类智能能自然理解视觉场景中丰富的语义维度,但现有语言-视觉概念学习方法仅限于颜色、形状等少数预定义基础轴,且多在合成数据集上进行。本文提出一种可扩展框架,能自适应识别真实场景中的图像相关概念轴,并在这些轴上实现视觉概念的语义锚定。该框架利用预训练的视觉-语言模型和通用提示策略,无需先验知识即可发现多样化概念轴;其通用概念编码器将视觉特征动态绑定到这些轴上,不引入额外参数。通过优化组合锚定目标,确保各轴可独立操控而不相互影响。在ImageNet、CelebA-HQ和AFHQ子集上的实验表明,该方法在复杂真实概念上展现出优越的编辑能力,具备强组合泛化性,优于现有视觉概念学习与文本编辑方法。代码已开源。
原文摘要 · Abstract (English)
Human intelligence effortlessly interprets visual scenes along a rich spectrum of semantic dimensions. However, existing approaches to language-grounded visual concept learning are limited to a few predefined primitive axes, such as color and shape, and are typically explored in synthetic datasets. In this work, we propose a scalable framework that adaptively identifies image-related concept axes and grounds visual concepts along these axes in real-world scenes. Leveraging a pretrained vision-language model and our universal prompting strategy, our framework identifies a diverse image-related axes without any prior knowledge. Our universal concept encoder adaptively binds visual features to the discovered axes without introducing additional model parameters for each concept. To ground visual concepts along the discovered axes, we optimize a compositional anchoring objective, which ensures that each axis can be independently manipulated without affecting others. We demonstrate the effectiveness of our framework on subsets of ImageNet, CelebA-HQ, and AFHQ, showcasing superior editing capabilities across diverse real-world concepts that are too varied to be manually predefined. Our method also exhibits strong compositional generalization, outperforming existing visual concept learning and text-based editing methods. The code is available at https://github.com/whieya/Language-grounded-VCL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。