arXiv:2509.00490cs.CVcs.AI2025-09

提出视频分组活动哈希技术,支持细粒度检索。

Multi-Focused Video Group Activities Hashing

  • 统一框架同时建模个体动作与群体互动
  • 在多个数据集上实现优异检索性能
  • 适合需要活动语义与视觉特征的场景

随着复杂场景中视频数据的爆炸式增长,快速检索群体活动成为迫切需求。然而,许多任务仅能针对整段视频进行检索,难以捕捉活动粒度。为此,我们首次提出一种新型时空交错视频哈希(STVH)技术。通过统一框架,STVH 同时建模个体对象动态与群体交互,捕获群体视觉特征与位置特征的时空演化。此外,在真实视频检索场景中,有时需关注活动特征,有时需关注对象视觉特征。为此,我们进一步提出增强版多聚焦时空视频哈希(M-STVH),通过层次化特征融合与多聚焦表征学习,使模型可联合关注活动语义特征与对象视觉特征。我们在公开数据集上进行了对比实验,STVH 与 M-STVH 均取得了优异结果。

原文摘要 · Abstract (English)

With the explosive growth of video data in various complex scenarios, quickly retrieving group activities has become an urgent problem. However, many tasks can only retrieve videos focusing on an entire video, not the activity granularity. To solve this problem, we propose a new STVH (spatiotemporal interleaved video hashing) technique for the first time. Through a unified framework, the STVH simultaneously models individual object dynamics and group interactions, capturing the spatiotemporal evolution on both group visual features and positional features. Moreover, in real-life video retrieval scenarios, it may sometimes require activity features, while at other times, it may require visual features of objects. We then further propose a novel M-STVH (multi-focused spatiotemporal video hashing) as an enhanced version to handle this difficult task. The advanced method incorporates hierarchical feature integration through multi-focused representation learning, allowing the model to jointly focus on activity semantics features and object visual features. We conducted comparative experiments on publicly available datasets, and both STVH and M-STVH can achieve excellent results.

视频检索哈希技术群体活动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。