arXiv:2410.15981cs.CVcs.LG2024-10被引 2

用知识图谱和合成图像提升模型在分布偏移下的图像分类能力

Robust Visual Representation Learning with Multi-modal Prior Knowledge for Image Classification Under Distribution Shift

  • 融合知识图谱与合成视觉元素,构建多模态先验引导表征学习
  • 在德国、中国、俄罗斯路标等数据集上准确率显著优于基线方法
  • 适合需要强泛化能力的跨域图像分类场景

尽管深度神经网络在计算机视觉中表现卓越,但在训练与测试数据分布发生变化时性能下降。本文提出知识引导视觉表征学习(KGV),一种基于分布的多模态先验学习方法,以增强分布偏移下的泛化能力。该方法融合两种模态:1)具有层级与关联关系的知识图谱(KG);2)由KG语义表示生成的合成视觉图像。从原始图像、合成图像及知识图谱中分别提取嵌入,并在统一潜在空间对齐。采用新型基于平移的知识图谱嵌入方法,将节点嵌入建模为高斯分布,关系嵌入建模为平移向量。实验表明,引入多模态先验可实现更正则化的表征学习,使模型在不同分布下具备更强泛化能力。在德国、中国、俄罗斯路标分类任务,mini-ImageNet及其变体,以及DVM-CAR数据集上均取得更高准确率与数据效率。

原文摘要 · Abstract (English)

Despite the remarkable success of deep neural networks (DNNs) in computer vision, they fail to remain high-performing when facing distribution shifts between training and testing data. In this paper, we propose Knowledge-Guided Visual representation learning (KGV) - a distribution-based learning approach leveraging multi-modal prior knowledge - to improve generalization under distribution shift. It integrates knowledge from two distinct modalities: 1) a knowledge graph (KG) with hierarchical and association relationships; and 2) generated synthetic images of visual elements semantically represented in the KG. The respective embeddings are generated from the given modalities in a common latent space, i.e., visual embeddings from original and synthetic images as well as knowledge graph embeddings (KGEs). These embeddings are aligned via a novel variant of translation-based KGE methods, where the node and relation embeddings of the KG are modeled as Gaussian distributions and translations, respectively. We claim that incorporating multi-model prior knowledge enables more regularized learning of image representations. Thus, the models are able to better generalize across different data distributions. We evaluate KGV on different image classification tasks with major or minor distribution shifts, namely road sign classification across datasets from Germany, China, and Russia, image classification with the mini-ImageNet dataset and its variants, as well as the DVM-CAR dataset. The results demonstrate that KGV consistently exhibits higher accuracy and data efficiency across all experiments.

视觉表征分布偏移知识图谱多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。