arXiv:2503.14524cs.CVcs.LG2025-03

只连接有时间关联的物体对,提升动态场景图生成精度

Salient Temporal Encoding for Dynamic Scene Graph Generation

  • 仅对时序相关物体重建时序边,构建稀疏显式时序关系
  • 在场景图检测上最高提升4.4%准确率,动作识别增0.6% mAP
  • 适合需要精准时序建模的视频理解任务

使用结构化时空场景图表示动态场景是一项新颖且极具挑战性的任务。关键在于学习物体间的时序交互关系,而不仅仅是空间关系。由于当前基准数据集缺乏显式标注的时序关系,现有方法通常在帧间所有物体对之间建立密集且抽象的时序连接。然而,并非所有时序连接都反映有意义的动态变化。本文提出一种新的时空场景图生成方法,仅在具有时序相关性的物体对之间构建时序连接,并将时序关系以显式边的形式表示在场景图中。这种稀疏且显式的时序表示使我们在场景图检测任务上相比强基线最高提升4.4%。此外,该方法还可用于提升下游视觉任务:应用于动作识别时,相较最先进方法取得0.6%的mAP提升。

原文摘要 · Abstract (English)

Representing a dynamic scene using a structured spatial-temporal scene graph is a novel and particularly challenging task. To tackle this task, it is crucial to learn the temporal interactions between objects in addition to their spatial relations. Due to the lack of explicitly annotated temporal relations in current benchmark datasets, most of the existing spatial-temporal scene graph generation methods build dense and abstract temporal connections among all objects across frames. However, not all temporal connections are encoding meaningful temporal dynamics. We propose a novel spatial-temporal scene graph generation method that selectively builds temporal connections only between temporal-relevant objects pairs and represents the temporal relations as explicit edges in the scene graph. The resulting sparse and explicit temporal representation allows us to improve upon strong scene graph generation baselines by up to $4.4\%$ in Scene Graph Detection. In addition, we show that our approach can be leveraged to improve downstream vision tasks. Particularly, applying our approach to action recognition, shows 0.6\% gain in mAP in comparison to the state-of-the-art

场景图生成时序建模视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。