构建首个含意图标注的多模态视频数据集,助力深度认知理解
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
- 采用分层描述框架,从事实到意图逐层解析视频内容
- 包含超2200万字描述,每样本平均225词,覆盖103000个视频样本
- 专为情感与意图识别等深层理解任务设计,适合多模态模型评估
本文提出VideoMind,一个面向深度视频认知理解的多模态视频数据集。该数据集包含10.3万个视频样本(其中3000个用于测试),每个视频均配有音频及系统化文本描述。每段描述分事实、抽象和意图三个层级,层层深入,总计超过2200万字,平均每样本约225词。其核心创新在于提供可反映上下文整合能力的意图表达,这些表达通过链式思维(COT)引导多模态大模型生成。所有描述均标注主体、地点、时间、事件、动作和意图,支持下游识别任务。我们建立3000个经人工验证的黄金标准样本作为基准评测集,并设计混合认知检索实验,使用多层级指标评估深度视频理解能力。公开发布模型(如InternVideo、VAST、UMT-L)在该数据集上的评测结果。数据已开源于GitHub、HuggingFace和OpenDataLab。
原文摘要 · Abstract (English)
This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired with audio and systematically detailed textual descriptions. Specifically, every video and its audio is described across three hierarchical layers (factual, abstract, and intent), progressing from surface to depth. It contains over 22 million words, averaging ~225 words per sample. VideoMind's key distinction from existing datasets is its provision of intent expressions, which require contextual integration across the entire video and are not directly observable. These deep-cognitive expressions are generated using a Chain-of-Thought (COT) approach, prompting the mLLM through step-by-step reasoning. Each description includes annotations for subject, place, time, event, action, and intent, supporting downstream recognition tasks. Crucially, we establish a gold-standard benchmark with 3,000 manually validated samples for evaluating deep-cognitive video understanding. We design hybrid-cognitive retrieval experiments, scored by multi-level retrieval metrics, to appropriately assess deep video comprehension. Evaluation results for models (e.g., InternVideo, VAST, UMT-L) are released. VideoMind serves as a powerful benchmark for fine-grained cross-modal alignment and advances fields requiring in-depth video understanding, such as emotion and intent recognition. The data is publicly available on GitHub, HuggingFace, and OpenDataLab, https://github.com/cdx-cindy/VideoMind.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。