arXiv:2504.14788cs.IR2025-04被引 7

聚焦多模态检索的高效表示学习,推动大模型落地应用

The 1st EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval

  • 组织首届多模态检索高效表示学习研讨会,聚焦大模型落地瓶颈
  • 呼吁建立效率评估指标与基准,破解训练部署效率难题
  • 适合关注大模型应用、信息检索效率的学术与产业研究者

多模态表示学习在人工智能领域受到广泛关注,主要得益于LLaMA、GPT、Mistral和CLIP等大型预训练多模态基础模型的成功。这些模型在网页搜索、跨模态检索、推荐系统等多种多模态信息检索(MIR)任务中表现出色。然而,由于参数量巨大,在将这些模型的表示能力应用于信息检索任务时,训练、部署和推理阶段均面临显著效率挑战,严重制约了基础模型在实际场景中的应用。为应对这一紧迫问题,我们提议在2025年万维网大会(Web Conference 2025)上举办首届EReL@MIR研讨会,邀请学术界与工业界参与者共同探讨新型解决方案、新兴问题、效率评估指标与基准。该研讨会旨在搭建交流平台,促进跨领域合作,推动大模型时代下多模态信息检索的高效且有效的表示学习发展。

原文摘要 · Abstract (English)

Multimodal representation learning has garnered significant attention in the AI community, largely due to the success of large pre-trained multimodal foundation models like LLaMA, GPT, Mistral, and CLIP. These models have achieved remarkable performance across various tasks of multimodal information retrieval (MIR), including web search, cross-modal retrieval, and recommender systems, etc. However, due to their enormous parameter sizes, significant efficiency challenges emerge across training, deployment, and inference stages when adapting these models' representation for IR tasks. These challenges present substantial obstacles to the practical adaptation of foundation models for representation learning in information retrieval tasks. To address these pressing issues, we propose organizing the first EReL@MIR workshop at the Web Conference 2025, inviting participants to explore novel solutions, emerging problems, challenges, efficiency evaluation metrics and benchmarks. This workshop aims to provide a platform for both academic and industry researchers to engage in discussions, share insights, and foster collaboration toward achieving efficient and effective representation learning for multimodal information retrieval in the era of large foundation models.

多模态检索高效学习大模型应用信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。