arXiv:2506.08817cs.CV2025-06被引 30

构建首个基于思维链的视频时空理解数据集,助力智能系统精准分析动态场景。

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

  • 采用思维链标注方式,生成19.2万条精细时空问答对
  • 模型在复杂时空任务上表现不佳,准确率普遍低于40%
  • 适合研究视频理解、多模态推理与智能交互的学者使用

视频内容理解对视频分析到交互系统等应用至关重要。尽管大规模视觉语言模型(VLMs)取得进展,但其在捕捉精细时空细节方面仍存在不足。为此,我们提出Video-CoT,一个基于思维链(CoT)方法的突破性数据集,包含192,000条细粒度时空问答对和23,000个高质量的CoT标注样本,为评估视频理解中的时空能力提供坚实基础。此外,我们构建了全面的评测基准,每个任务包含750张图像并配备定制化评估指标。大量实验表明,当前VLMs在实现满意性能方面面临显著挑战,凸显了有效时空理解的难度。整体而言,Video-CoT数据集与基准为多媒体理解研究开辟新路径,并支持需要先进视频分析能力的智能系统未来发展。项目主页:https://video-cot.github.io/

原文摘要 · Abstract (English)

Video content comprehension is essential for various applications, ranging from video analysis to interactive systems. Despite advancements in large-scale vision-language models (VLMs), these models often struggle to capture the nuanced, spatiotemporal details essential for thorough video analysis. To address this gap, we introduce Video-CoT, a groundbreaking dataset designed to enhance spatiotemporal understanding using Chain-of-Thought (CoT) methodologies. Video-CoT contains 192,000 fine-grained spa-tiotemporal question-answer pairs and 23,000 high-quality CoT-annotated samples, providing a solid foundation for evaluating spatiotemporal understanding in video comprehension. Additionally, we provide a comprehensive benchmark for assessing these tasks, with each task featuring 750 images and tailored evaluation metrics. Our extensive experiments reveal that current VLMs face significant challenges in achieving satisfactory performance, high-lighting the difficulties of effective spatiotemporal understanding. Overall, the Video-CoT dataset and benchmark open new avenues for research in multimedia understanding and support future innovations in intelligent systems requiring advanced video analysis capabilities. By making these resources publicly available, we aim to encourage further exploration in this critical area. Project website:https://video-cot.github.io/ .

视频理解思维链多模态数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。