arXiv:2512.18448cs.CV2025-12AAAI被引 2

提出基于物体中心的视频片段定位框架,提升对特定物体交互的精准定位能力。

Object-Centric Framework for Video Moment Retrieval

  • 通过场景图解析提取查询相关物体,构建物体级特征序列
  • 在三个数据集上均超越现有最佳方法,最高提升6.2% mAP
  • 适合需要细粒度物体推理的任务,如复杂交互理解

现有视频片段定位方法依赖帧或片段级别的全局视觉与语义特征,难以捕捉细粒度物体语义与外观,尤其在涉及特定实体及其交互的对象导向查询中表现不足。本文提出一种新型物体中心框架,首先使用场景图解析器提取查询相关物体,并从视频帧生成场景图以表征物体及其关系;基于场景图构建编码丰富视觉与语义信息的物体级特征序列,再通过关系轨迹变换器建模物体间的时空关联。该框架显式捕捉物体状态随时间的变化,实现更精确的片段定位。在Charades-STA、QVHighlights和TACoS三个基准上评估,结果表明本方法全面优于现有最先进方法。

原文摘要 · Abstract (English)

Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks.

视频定位物体中心场景图时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。