通过解耦视觉特征提升模型对新组合的泛化能力
Independent Density Estimation
- 用解耦视觉表示建模词与图像特征的独立关系
- 在未见组合上优于现有模型,跨数据集表现更优
- 适合研究视觉语言模型泛化与可解释性的人参考
大规模视觉-语言模型在图像描述和条件图像生成等任务中取得了显著成果,但仍难以实现类人水平的组合泛化。本文提出一种名为独立密度估计(IDE)的新方法,旨在学习句子中单个词汇与图像对应特征之间的关联,从而实现组合泛化。基于IDE理念,我们构建了两个模型:第一个使用完全解耦的视觉表示作为输入,第二个则利用变分自编码器从原始图像中获取部分解耦特征。此外,我们提出一种基于熵的组合推理方法,用于融合句子中每个词的预测结果。在多个数据集上的评估表明,我们的模型在未见组合上的泛化能力显著优于当前主流模型。
原文摘要 · Abstract (English)
Large-scale Vision-Language models have achieved remarkable results in various domains, such as image captioning and conditioned image generation. Nevertheless, these models still encounter difficulties in achieving human-like compositional generalization. In this study, we propose a new method called Independent Density Estimation (IDE) to tackle this challenge. IDE aims to learn the connection between individual words in a sentence and the corresponding features in an image, enabling compositional generalization. We build two models based on the philosophy of IDE. The first one utilizes fully disentangled visual representations as input, and the second leverages a Variational Auto-Encoder to obtain partially disentangled features from raw images. Additionally, we propose an entropy-based compositional inference method to combine predictions of each word in the sentence. Our models exhibit superior generalization to unseen compositions compared to current models when evaluated on various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。