构建多视角多领域多模态检索基准,突破现有模型单一视角局限。
MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval

- 提出多视角标注的10.5万条开放性查询数据集,覆盖多领域多模态
- 新方法SPIN通过噪声向量引导,使检索结果覆盖更广视角,提升37%覆盖度
- 适合关注开放性检索、多模态理解与公平性评估的研究者
信息检索日益面向可多角度回答的开放性问题。现有基准多聚焦封闭式问题,即便开放性基准也大多仅涵盖单一主题领域和模态。本文提出Multi³IR,一个评估检索器在跨多领域、多模态下覆盖开放性问题多元视角能力的基准。该数据集包含104.9万条Stack Exchange查询,每条均附有捕捉其隐含视角的描述标注。我们进一步提出SPIN方法,通过学习噪声向量有效引导嵌入朝多样化且语义合理的方向演化。实验表明,现有多模态检索器普遍存在单视角偏见;而SPIN在Multi³IR上显著提升视角覆盖率,并在未见的开放性检索基准上具有良好泛化能力。数据集与代码已开源。
原文摘要 · Abstract (English)
Information retrieval (IR) increasingly targets open-ended queries that admit diverse perspectives. Existing IR benchmarks, however, focus primarily on closed-ended queries, while even open-ended benchmarks largely consist of queries whose supporting documents span a single subject domain and modality. We introduce Multi$^3$IR, a benchmark that evaluates how well retrievers cover the multifaceted perspectives of open-ended queries across diverse domains and modalities. It comprises 104.9K Stack Exchange queries, each annotated with perspective descriptions that capture the query's implicit viewpoints. We further propose SPIN, a parameter- and label-efficient method that learns noise vectors to steer embeddings toward diverse yet meaningful semantic directions. Experiments show that existing multimodal retrievers suffer from single-perspective bias, while SPIN substantially improves perspective coverage on Multi$^3$IR and generalizes well to unseen open-ended IR benchmarks. The dataset and experimental code are available at https://github.com/seokwon99/Multi3IR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。