提出视频语言理解新评测任务,检验模型长时序记忆与时间定位能力。
NeMo: Needle in a Montage for Video-Language Understanding
- 设计'蒙太奇中的针'任务,评估视频模型对长时间上下文的回忆与定位能力。
- 构建包含31,378个问答对的NeMoBench基准,覆盖13,486段时长从秒到小时不等的视频。
- 提供自动化数据生成流水线,支持持续更新,适合评估最新视频大模型性能。
视频大语言模型(VideoLLMs)的进展亟需新的评估协议与基准。受大语言模型中经典的‘针在草堆’测试启发,我们提出一项名为‘蒙太奇中的针’(NeMo)的新任务,用于评估先进VideoLLMs的时间理解能力。该任务聚焦于两种关键能力:检索式长上下文回忆与时间定位。为生成该任务所需的视频问答数据,我们开发了一个可扩展的自动化数据生成流水线,实现高质量数据合成。基于此流水线,我们构建了以该任务为核心的视频-语言基准NeMoBench。完整版NeMoBench包含来自13,486段视频的31,378个自动生成的问答对,视频时长跨度从秒级至小时级。实验表明,该流水线能可靠、自动地生成高质量评估数据,使NeMoBench可随最新视频持续更新。我们在该基准上评估了20个最先进模型,提供了详尽结果与关键洞察,揭示其能力与局限性。项目页面见:https://lavi-lab.github.io/NeMoBench。
原文摘要 · Abstract (English)
Recent advances in video large language models (VideoLLMs) call for new evaluation protocols and benchmarks for video-language understanding. Inspired by the needle in a haystack test widely used by LLMs, we introduce a novel task of Needle in a Montage (NeMo), which is designed to assess the temporal understanding capabilities of advanced VideoLLMs. Specifically, the proposed task focuses on two fundamental abilities critical for temporal understanding, i.e., retrieval-style long-context recall and temporal grounding. To generate video question answering data for our task, we develop a scalable automated data generation pipeline that facilitates high-quality data synthesis. Built upon the proposed pipeline, we present NeMoBench, a video-language benchmark centered on our task. Specifically, our full set of NeMoBench features 31,378 automatically generated question-answer (QA) pairs from 13,486 videos with various durations ranging from seconds to hours. Experiments demonstrate that our pipeline can reliably and automatically generate high-quality evaluation data, enabling NeMoBench to be continuously updated with the latest videos. We evaluate 20 state-of-the-art models on our benchmark, providing extensive results and key insights into their capabilities and limitations. Our project page is available at: https://lavi-lab.github.io/NeMoBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。