arXiv:2601.10592cs.CV2026-01被引 11

构建1亿级动作视频数据集,推动机器理解物理世界行为

Action100M: A Large-scale Video Action Dataset

论文配图:Action100M: A Large-scale Video Action Dataset
图 1 · 摘自论文原文
  • 用自动生成管道从120万条教学视频提取时序动作片段
  • 产出超1亿个带丰富描述的动作段落,支持开放词汇识别
  • 适合视频理解、世界建模研究者,助力零样本模型训练

从120万条互联网教学视频(总计14.6年时长)中构建大规模视频动作数据集Action100M,生成约1亿个时间定位的动作片段,具备开放词汇动作标注和丰富描述。该数据集通过全自动流程实现:(i) 基于V-JEPA 2嵌入的分层时序分割;(ii) 生成多层级帧与片段描述,形成树状描述结构(Tree-of-Captions);(iii) 利用推理模型GPT-OSS-120B在多轮Self-Refine机制下聚合证据,输出结构化标注(简明/详细动作、执行者、简明/详细描述)。在Action100M上训练的VL-JEPA展现持续的数据规模提升效果,并在多个动作识别基准上实现优异零样本性能,确立其作为视频理解与世界建模可扩展研究新基线的地位。

原文摘要 · Abstract (English)

Inferring physical actions from visual observations is a fundamental capability for advancing machine intelligence in the physical world. Achieving this requires large-scale, open-vocabulary video action datasets that span broad domains. We introduce Action100M, a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding O(100 million) temporally localized segments with open-vocabulary action supervision and rich captions. Action100M is generated by a fully automated pipeline that (i) performs hierarchical temporal segmentation using V-JEPA 2 embeddings, (ii) produces multi-level frame and segment captions organized as a Tree-of-Captions, and (iii) aggregates evidence with a reasoning model (GPT-OSS-120B) under a multi-round Self-Refine procedure to output structured annotations (brief/detailed action, actor, brief/detailed caption). Training VL-JEPA on Action100M demonstrates consistent data-scaling improvements and strong zero-shot performance across diverse action recognition benchmarks, establishing Action100M as a new foundation for scalable research in video understanding and world modeling.

视频理解动作识别大规模数据集自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。