arXiv:2510.13276cs.CVcs.CL2025-10被引 1

评测视觉语言模型在长上下文中的信息保持能力。

MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models

  • 设计跨模态长上下文评测基准,覆盖图文视频多类型数据。
  • 8项任务测试不同长度上下文下模型的准确性表现。
  • 揭示大模型在长上下文场景中信息遗忘问题,适合研究者参考。

大型视觉语言模型(LVLMs)的上下文窗口持续扩展,但更长的上下文并不意味着能有效利用。当前对长上下文忠实度的评估主要集中在纯文本领域,多模态评估仍局限于短上下文。为此,我们提出MMLongCite,一个全面的基准,用于评估LVLM在长上下文场景下的信息忠实度。该基准包含8个任务,覆盖6个不同的上下文长度区间,融合文本、图像和视频等多种模态。对主流LVLM的评估显示,其在处理长多模态上下文时忠实度有限。此外,我们深入分析了上下文长度及关键内容位置对模型表现的影响。

原文摘要 · Abstract (English)

The rapid advancement of large vision language models (LVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness are predominantly focused on the text-only domain, while multimodal assessments remain limited to short contexts. To bridge this gap, we introduce MMLongCite, a comprehensive benchmark designed to evaluate the fidelity of LVLMs in long-context scenarios. MMLongCite comprises 8 distinct tasks spanning 6 context length intervals and incorporates diverse modalities, including text, images, and videos. Our evaluation of state-of-the-art LVLMs reveals their limited faithfulness in handling long multimodal contexts. Furthermore, we provide an in-depth analysis of how context length and the position of crucial content affect the faithfulness of these models.

视觉语言模型长上下文多模态评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。