arXiv:2601.14714cs.IR2026-01

统一图文与多语言检索,用共享编码器提升效率与准确率

Unified Multimodal and Multilingual Retrieval via Multi-Task Learning with NLU Integration

  • 通过多任务学习融合图像、文本与意图理解的特征表示
  • 在多语言图文检索中显著优于专用文本模型,且减少存储开销
  • 首次联合优化多语言检索与自然语言理解,适合跨模态应用

多模态检索系统通常使用视觉语言模型(VLMs)将图像和文本独立编码到共享嵌入空间。尽管引入了文本编码器,VLMs 在纯文本检索任务上仍持续落后于专用文本模型。此外,增加额外文本编码器会带来存储与推理开销上升,尤其在多语言场景下加剧检索效率问题。为解决上述局限,我们提出一种多任务学习框架,统一图像、长文本、短文本及意图丰富查询的特征表示。据我们所知,这是首个在单一框架中联合优化多语言图像检索、文本检索与自然语言理解(NLU)任务的工作。该方法通过共享文本编码器并融入NLU特征,增强意图理解能力与检索准确性。

原文摘要 · Abstract (English)

Multimodal retrieval systems typically employ Vision Language Models (VLMs) that encode images and text independently into vectors within a shared embedding space. Despite incorporating text encoders, VLMs consistently underperform specialized text models on text-only retrieval tasks. Moreover, introducing additional text encoders increases storage, inference overhead, and exacerbates retrieval inefficiencies, especially in multilingual settings. To address these limitations, we propose a multi-task learning framework that unifies the feature representation across images, long and short texts, and intent-rich queries. To our knowledge, this is the first work to jointly optimize multilingual image retrieval, text retrieval, and natural language understanding (NLU) tasks within a single framework. Our approach integrates image and text retrieval with a shared text encoder that is enhanced by NLU features for intent understanding and retrieval accuracy.

多模态检索多语言NLU多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。