arXiv:2509.12145cs.CV2025-09ICCV被引 5

用大模型实现视频流中动作层级识别与自由描述生成。

Open-ended Hierarchical Streaming Video Understanding with Vision Language Models

论文配图:Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
图 1 · 摘自论文原文
  • 利用大模型将细粒度动作归类为高层事件,填补标注数据不足。
  • 提出OpenHOUSE系统,在相邻动作边界检测上性能接近翻倍提升。
  • 适合关注视频理解与生成融合的科研人员与工程师。

我们提出了层次化视频流理解任务,结合在线时间动作定位与自由形式描述生成。由于缺乏具有层次性和细粒度时间标注的数据集,我们证明大语言模型可有效将原子动作聚类为更高层级事件,从而丰富现有数据集。随后,我们提出OpenHOUSE(面向事件的开放性层次在线理解系统),将视频流动作感知从动作分类拓展至更复杂场景。该系统配备专用流式模块,能精准检测紧密相邻动作之间的边界,其性能几乎达到现有方法直接扩展结果的两倍。我们展望未来视频流感知将融合强大生成模型,OpenHOUSE为此方向迈出关键一步。

原文摘要 · Abstract (English)

We introduce Hierarchical Streaming Video Understanding, a task that combines online temporal action localization with free-form description generation. Given the scarcity of datasets with hierarchical and fine-grained temporal annotations, we demonstrate that LLMs can effectively group atomic actions into higher-level events, enriching existing datasets. We then propose OpenHOUSE (Open-ended Hierarchical Online Understanding System for Events), which extends streaming action perception beyond action classification. OpenHOUSE features a specialized streaming module that accurately detects boundaries between closely adjacent actions, nearly doubling the performance of direct extensions of existing methods. We envision the future of streaming action perception in the integration of powerful generative models, with OpenHOUSE representing a key step in that direction.

视频理解大模型流式处理动作识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。