统一图文与多语言检索,用共享编码器提升效率与准确率
Unified Multimodal and Multilingual Retrieval via Multi-Task Learning with NLU Integration
- 通过多任务学习融合图像、文本与意图理解的特征表示
- 在多语言图文检索中显著优于专用文本模型,且减少存储开销
- 首次联合优化多语言检索与自然语言理解,适合跨模态应用
多模态检索系统通常使用视觉语言模型(VLMs)将图像和文本独立编码到共享嵌入空间。尽管引入了文本编码器,VLMs 在纯文本检索任务上仍持续落后于专用文本模型。此外,增加额外文本编码器会带来存储与推理开销上升,尤其在多语言场景下加剧检索效率问题。为解决上述局限,我们提出一种多任务学习框架,统一图像、长文本、短文本及意图丰富查询的特征表示。据我们所知,这是首个在单一框架中联合优化多语言图像检索、文本检索与自然语言理解(NLU)任务的工作。该方法通过共享文本编码器并融入NLU特征,增强意图理解能力与检索准确性。
原文摘要 · Abstract (English)
Multimodal retrieval systems typically employ Vision Language Models (VLMs) that encode images and text independently into vectors within a shared embedding space. Despite incorporating text encoders, VLMs consistently underperform specialized text models on text-only retrieval tasks. Moreover, introducing additional text encoders increases storage, inference overhead, and exacerbates retrieval inefficiencies, especially in multilingual settings. To address these limitations, we propose a multi-task learning framework that unifies the feature representation across images, long and short texts, and intent-rich queries. To our knowledge, this is the first work to jointly optimize multilingual image retrieval, text retrieval, and natural language understanding (NLU) tasks within a single framework. Our approach integrates image and text retrieval with a shared text encoder that is enhanced by NLU features for intent understanding and retrieval accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。