arXiv:2605.26941cs.IRcs.MM2026-05中稿 · as a workshop prop…

聚焦大模型时代多模态检索的高效表示学习挑战

The 2nd EReL@MIR Workshop on Efficient Representation Learning for Multimodal Information Retrieval

  • 组织跨学术与产业界的研讨会,共议效率问题
  • 针对大模型在检索任务中训练部署瓶颈提出新解法
  • 适合关注多模态检索与模型效率的研究者

多模态表示学习在人工智能领域受到越来越多关注,主要得益于Qwen、LLaVA和CLIP等大型预训练多模态基础模型的出色表现。这些模型在网页搜索、跨模态检索和推荐系统等多种多模态信息检索(MIR)任务上表现出色。然而,其庞大的参数量在训练、部署和推理过程中带来显著的效率瓶颈,制约了基础模型在信息检索中的实际应用。为应对这一挑战,我们计划于MM 2026举办EReL@MIR研讨会,汇聚学术界与工业界研究人员,探讨新兴解决方案、开放挑战以及基础模型时代多模态IR表示学习的新效率度量与基准。研讨会官网:https://erel-mir.github.io/。

原文摘要 · Abstract (English)

Multimodal representation learning has attracted increasing attention in AI, driven by the strong performance of large, pretrained multimodal foundation models such as Qwen, LLaVA, and CLIP. These models deliver impressive performance on a range of multimodal information retrieval (MIR) tasks, including web search, cross-modal retrieval, and recommender systems. Yet their massive parameter counts create major efficiency bottlenecks when adapting their representations for IR tasks during training, deployment, and inference. These limitations hinder the practical use of foundation models for representation learning in information retrieval. To address these issues, we propose organizing the EReL@MIR workshop at MM 2026, bringing together researchers from academia and industry to discuss emerging solutions, open challenges, and new efficiency metrics and benchmarks for multimodal IR representation learning in the foundation-model era. The workshop's official website is available at https://erel-mir.github.io/.

多模态检索高效学习基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。