首个面向流式视频问答的时序推理数据集,支持动态答案与多模态思维链分析。
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
- 构建每秒级动态标注架构,捕捉视频中答案随时间演化的特性。
- 引入显式思维链生成机制,实现基于关键帧对齐与状态演变的逻辑推理路径。
- 适合研究流媒体理解、复杂时序推理及可解释多模态模型的开发者和学者。
流媒体视频应用的快速发展要求多模态模型具备更强的时序动态理解与复杂推理能力。然而,现有视频问答(VideoQA)数据集存在两大缺陷:1)静态标注无法反映答案在时序视频流中的动态变化;2)缺乏明确的推理过程标注,限制了模型的可解释性与逻辑推断能力。为此,我们提出StreamingCoT,首个专为流式视频问答中时序演化推理与多模态思维链(CoT)任务设计的数据集。框架首先建立动态分层标注体系,通过相似性融合生成每秒密集描述,并构建时序依赖的语义片段,同时约束问题-答案对符合时序演化模式。进一步提出显式思维链生成范式:利用关键帧语义对齐提取时空对象,借助大语言模型推导基于对象状态转移的推理路径,并通过人工验证确保逻辑一致性。该数据集为推进流式视频理解、复杂时序推理与多模态推理研究奠定基础。StreamingCoT及其构建工具包可访问 https://github.com/Fleeting-hyh/StreamingCoT。
原文摘要 · Abstract (English)
The rapid growth of streaming video applications demands multimodal models with enhanced capabilities for temporal dynamics understanding and complex reasoning. However, current Video Question Answering (VideoQA) datasets suffer from two critical limitations: 1) Static annotation mechanisms fail to capture the evolving nature of answers in temporal video streams, and 2) The absence of explicit reasoning process annotations restricts model interpretability and logical deduction capabilities. To address these challenges, We introduce StreamingCoT, the first dataset explicitly designed for temporally evolving reasoning in streaming VideoQA and multimodal Chain-of-Thought (CoT) tasks. Our framework first establishes a dynamic hierarchical annotation architecture that generates per-second dense descriptions and constructs temporally-dependent semantic segments through similarity fusion, paired with question-answer sets constrained by temporal evolution patterns. We further propose an explicit reasoning chain generation paradigm that extracts spatiotemporal objects via keyframe semantic alignment, derives object state transition-based reasoning paths using large language models, and ensures logical coherence through human-verified validation. This dataset establishes a foundation for advancing research in streaming video understanding, complex temporal reasoning, and multimodal inference. Our StreamingCoT and its construction toolkit can be accessed at https://github.com/Fleeting-hyh/StreamingCoT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。