无需标注,模型自动从视频流中识别出嵌套的事件结构。
Generalized Event Partonomy Inference with Structured Hierarchical Predictive Learning
- 用分层递归预测器建模不同时间尺度的视频动态。
- 在三个数据集上达到顶尖流式处理效果,事件边界精准匹配人类感知。
- 适合研究视频理解、时序建模与人类认知对齐的学者。
人类自然将连续体验视为嵌套的层次化事件结构,细粒度动作嵌套于更宏观的例行程序中。计算机视觉中复现这一结构需模型能不仅事后分割视频,更具备预测性与层次性。我们提出PARS,一种统一框架,可直接从视频流中无监督学习多尺度事件结构。PARS将感知组织为分层递归预测器,各层级以自身时间粒度运行:低层建模短期动态,高层通过基于注意力的反馈整合长期上下文。事件边界自然表现为预测误差的瞬时峰值,生成时间连贯、嵌套分明的结构,符合人类事件感知中的包含关系。在Breakfast Actions、50 Salads和Assembly 101三个基准上评估,PARSE在流式方法中表现最优,并在时间对齐(H-GEBD)与结构一致性(TED, hF1)上媲美离线基线。结果表明,在不确定性下进行预测学习,是实现类人时间抽象与组合事件理解的可扩展路径。
原文摘要 · Abstract (English)
Humans naturally perceive continuous experience as a hierarchy of temporally nested events, fine-grained actions embedded within coarser routines. Replicating this structure in computer vision requires models that can segment video not just retrospectively, but predictively and hierarchically. We introduce PARSE, a unified framework that learns multiscale event structure directly from streaming video without supervision. PARSE organizes perception into a hierarchy of recurrent predictors, each operating at its own temporal granularity: lower layers model short-term dynamics while higher layers integrate longer-term context through attention-based feedback. Event boundaries emerge naturally as transient peaks in prediction error, yielding temporally coherent, nested partonomies that mirror the containment relations observed in human event perception. Evaluated across three benchmarks, Breakfast Actions, 50 Salads, and Assembly 101, PARSE achieves state-of-the-art performance among streaming methods and rivals offline baselines in both temporal alignment (H-GEBD) and structural consistency (TED, hF1). The results demonstrate that predictive learning under uncertainty provides a scalable path toward human-like temporal abstraction and compositional event understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。