arXiv:2506.03144cs.CVcs.CL2025-06NeurIPS被引 5

首个多语言交叉条件语义检索数据集,解决跨模态理解难题

MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query

  • 构建5语言32万查询的多条件交叉检索数据集
  • 新框架Coral提升45.9%性能,兼顾细粒度条件与全局语义
  • 适合多语言视觉搜索、跨模态理解研究者使用

语义检索对现代应用至关重要,但现有研究受限于单语言、单图像或单一检索条件。实际场景常涉及多图像、多条件交叉查询。本文提出MERIT,首个支持多语言的交叉条件语义检索数据集,包含32万条查询、13.5万件商品,覆盖5种语言和7类商品。实验发现现有模型仅关注全局语义,忽略查询中的具体条件。为此,我们提出Coral微调框架,通过嵌入重建保留细粒度条件信息,并结合对比学习提取完整全局语义。在MERIT上,Coral相比传统方法性能提升45.9%,并在8个主流检索基准上验证了强泛化能力。本工作贡献包括新数据集、关键问题识别及创新微调框架,为交叉条件语义检索奠定基础。

原文摘要 · Abstract (English)

Semantic retrieval is crucial for modern applications yet remains underexplored in current research. Existing datasets are limited to single languages, single images, or singular retrieval conditions, often failing to fully exploit the expressive capacity of visual information as evidenced by maintained performance when images are replaced with captions. However, practical retrieval scenarios frequently involve interleaved multi-condition queries with multiple images. Hence, this paper introduces MERIT, the first multilingual dataset for interleaved multi-condition semantic retrieval, comprising 320,000 queries with 135,000 products in 5 languages, covering 7 distinct product categories. Extensive experiments on MERIT identify existing models's limitation: focusing solely on global semantic information while neglecting specific conditional elements in queries. Consequently, we propose Coral, a novel fine-tuning framework that adapts pre-trained MLLMs by integrating embedding reconstruction to preserve fine-grained conditional elements and contrastive learning to extract comprehensive global semantics. Experiments demonstrate that Coral achieves a 45.9% performance improvement over conventional approaches on MERIT, with strong generalization capabilities validated across 8 established retrieval benchmarks. Collectively, our contributions - a novel dataset, identification of critical limitations in existing approaches, and an innovative fine-tuning framework - establish a foundation for future research in interleaved multi-condition semantic retrieval.

语义检索多语言跨模态Mixture of Conditions

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。