构建首个评估大模型时间事件理解能力的基准,挑战现有模型表现。
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
- 设计包含822个视频-文本问题的RTime-QA基准,聚焦原子时间事件理解。
- SOTA模型Qwen2-VL在严格准确率上仅达34.6,远低于人类水平。
- 引入1.4万条指令微调数据RTime-IT,显著提升模型时间理解能力。
准确理解原子时间事件对视频理解至关重要。然而,现有视频-语言基准往往难以有效评估大模型(LMMs)的时间事件理解能力,因其可被图像-语言模型解决。本文提出RTime-QA,一个专为评估LMMs原子时间事件理解能力而设计的新基准。该基准包含822个高质量、人工精心标注的视频-文本问题,每个问题对应一个描绘原子时间事件的视频,并配有正确答案及时间否定描述,以精准测试时间理解能力。为推动模型发展,我们进一步构建了1.4万条指令微调数据集RTime-IT,采用与RTime-QA一致的标注流程。大量实验表明,RTime-QA对现有LMMs构成严峻挑战:SOTA模型Qwen2-VL在严格准确率(strict-ACC)上仅为34.6,显著落后于人类表现。同时,通过在RTime-IT上微调,Qwen2-VL在RTime-QA上的得分提升至65.9,验证了数据集的有效性。
原文摘要 · Abstract (English)
Understanding accurate atomic temporal event is essential for video comprehension. However, current video-language benchmarks often fall short to evaluate Large Multi-modal Models' (LMMs) temporal event understanding capabilities, as they can be effectively addressed using image-language models. In this paper, we introduce RTime-QA, a novel benchmark specifically designed to assess the atomic temporal event understanding ability of LMMs. RTime-QA comprises 822 high-quality, carefully-curated video-text questions, each meticulously annotated by human experts. Each question features a video depicting an atomic temporal event, paired with both correct answers and temporal negative descriptions, specifically designed to evaluate temporal understanding. To advance LMMs' temporal event understanding ability, we further introduce RTime-IT, a 14k instruction-tuning dataset that employs a similar annotation process as RTime-QA. Extensive experimental analysis demonstrates that RTime-QA presents a significant challenge for LMMs: the state-of-the-art model Qwen2-VL achieves only 34.6 on strict-ACC metric, substantially lagging behind human performance. Furthermore, our experiments reveal that RTime-IT effectively enhance LMMs' capacity in temporal understanding. By fine-tuning on RTime-IT, our Qwen2-VL achieves 65.9 on RTime-QA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。