arXiv:2412.13510cs.CVcs.CL2024-12中稿 · AAAI被引 3

动态适配器让低资源语言图文检索更灵活精准

Dynamic Adapter with Semantics Disentangling for Cross-lingual Cross-modal Retrieval

  • 根据输入描述特征动态生成适配器参数
  • 在多数据集上提升低资源语言检索效果
  • 适合需要跨语言图文匹配的轻量级应用

现有跨模态检索方法依赖大规模视觉-语言配对数据,难以高效构建低资源语言的检索模型。为此,跨语言跨模态检索(CCR)关注在不使用目标语言人工标注数据的情况下,实现视觉与低资源语言间的对齐。当前主流方法采用适配器模块,将视觉-语言预训练(VLP)模型在源语言上的对齐能力迁移至目标语言,但传统适配器参数固定,难以应对目标语言描述表达多样性的挑战。为此,本文提出动态适配器与语义解耦方法(DASD),其参数根据输入描述特征动态生成。考虑到输入描述的语义与表达风格共同影响编码方式,我们设计语义解耦模块,分离出语义相关与语义无关特征,使生成的适配器更契合输入描述特性。在两个图像-文本数据集和一个视频-文本数据集上的大量实验表明,该方法显著提升了跨语言跨模态检索性能,并展现出对多种VLP模型的良好兼容性。

原文摘要 · Abstract (English)

Existing cross-modal retrieval methods typically rely on large-scale vision-language pair data. This makes it challenging to efficiently develop a cross-modal retrieval model for under-resourced languages of interest. Therefore, Cross-lingual Cross-modal Retrieval (CCR), which aims to align vision and the low-resource language (the target language) without using any human-labeled target-language data, has gained increasing attention. As a general parameter-efficient way, a common solution is to utilize adapter modules to transfer the vision-language alignment ability of Vision-Language Pretraining (VLP) models from a source language to a target language. However, these adapters are usually static once learned, making it difficult to adapt to target-language captions with varied expressions. To alleviate it, we propose Dynamic Adapter with Semantics Disentangling (DASD), whose parameters are dynamically generated conditioned on the characteristics of the input captions. Considering that the semantics and expression styles of the input caption largely influence how to encode it, we propose a semantic disentangling module to extract the semantic-related and semantic-agnostic features from the input, ensuring that generated adapters are well-suited to the characteristics of input caption. Extensive experiments on two image-text datasets and one video-text dataset demonstrate the effectiveness of our model for cross-lingual cross-modal retrieval, as well as its good compatibility with various VLP models.

跨模态检索低资源语言动态适配器语义解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。