首个针对长上下文多模态模型的评测基准,揭示现有模型在理解长序列时的深层缺陷。
MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios

- 构建涵盖文本、文档、视频的四类长上下文检索任务,系统评估多模态模型能力。
- 发现主流模型依赖表层特征匹配,难以捕捉深层语义与结构关联,性能随上下文长度下降明显。
- 揭示不同模态对冗余信息的鲁棒性差异,为模型优化提供关键线索。
近期进展显著扩展了多模态嵌入模型(MEMs)的理论上下文窗口,但更大的上下文窗口并不必然带来对长上下文多模态输入的有效理解与表征,这仍是实际部署的关键瓶颈。为解决该场景下缺乏系统评估的问题,我们提出 MMLongEmbed,这是首个面向长上下文场景的多模态嵌入模型综合评测基准。MMLongEmbed 包含覆盖多个上下文长度范围的四类检索任务,涵盖文本、文档和视频模态。通过对最先进模型的广泛评估,我们发现当前架构严重依赖表层特征匹配,难以捕获深层语义与结构依赖。此外,性能退化随上下文长度和关键信息位置呈现系统性变化,且不同模态对冗余上下文信息表现出显著不同的鲁棒性。为保证可复现性,该基准及代码已公开。
原文摘要 · Abstract (English)
Recent advancements have significantly expanded the theoretical context windows of Multimodal Embedding Models (MEMs). However, larger context windows do not necessarily translate into effective comprehension and representation of long-context multimodal inputs, which remains a critical bottleneck for real-world deployment. To address the lack of systematic evaluation in this setting, we introduce MMLongEmbed, the first comprehensive benchmark for evaluating MEMs in long-context scenarios. MMLongEmbed comprises four retrieval tasks spanning multiple context-length ranges, covering text, document, and video modalities. Through extensive evaluation of state-of-the-art models, we find that current architectures rely heavily on superficial feature matching and struggle to capture deep semantic and structural dependencies. We further observe that performance degradation varies systematically with context length and key information placement. Moreover, models exhibit substantially different robustness to redundant contextual information across modalities. For reproducibility, the benchmark and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。