arXiv:2609.08755cs.CVcs.AI2026-09

构建时序精细的视频语言数据集,支持长期动态建模

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

论文配图:Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
图 1 · 摘自论文原文
  • 提供长达10至30分钟的视频,标注时间分辨率高
  • 涵盖动作、实体、交互和上下文线索的细粒度时序信息
  • 适合长程建模、指令生成与视觉动态表示学习

许多新兴的视频语言建模任务需要系统超越片段级抽象,对随时间展开的视觉内容进行建模。然而,现有视频数据集大多依赖粗粒度或稀疏对齐的监督信号,压缩了时间变化,限制了模型学习连续视觉动态可复用表征的能力。我们提出Kairos,一个用于视频-语言建模的时序解析数据集。Kairos包含持续时间从十分钟到半小时的长视频,配有细粒度的时间对齐标注。标注内容涵盖持续动作、实体出现与属性、交互行为以及随时间演化的上下文线索。该时序解析结构支持细粒度评估、长程建模与推理、指令数据构建、表征学习及视频生成。Kairos为建模时间维度上的视觉体验提供了通用基础。

原文摘要 · Abstract (English)

Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most existing video datasets rely on coarse or sparsely aligned supervision, which compresses temporal variation and limits the ability of models to learn reusable representations of continuous visual dynamics. We introduce Kairos, a video dataset for video-language modeling with time-resolved annotations. Kairos consists of long-duration videos, ranging from ten minutes to half an hour, annotated with fine-grained temporal alignment. The annotations capture ongoing actions, entity appearances and attributes, interactions, and evolving contextual cues along the video timeline. This time-resolved structure supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation. Kairos provides a general-purpose foundation for modeling visual experiences over time.

视频语言时序建模数据集长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。