提出时空协同建模方法,提升航拍视频场景图的准确性与连贯性。
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage
- 分层聚合+循环时序优化,同步捕捉多尺度空间与长程时序关系。
- 在ASPIRe和AeroEye-v1.0上显著超越现有方法,准确率提升显著。
- 专为航拍设计的新数据集含五类交互,适合动态场景理解研究者。
自动驾驶、监控和体育分析等应用中视频数据的快速增长,对动态场景理解提出了更高要求。尽管静态场景图生成和早期视频场景图生成已取得进展,但现有方法常因表示碎片化,难以同时捕捉精细的空间细节与长程时序依赖。为此,本文提出时空分层循环场景图(THYME)方法,通过分层特征聚合与循环时序优化的协同机制,有效建模多尺度空间上下文并强化帧间时序一致性,生成更准确且连贯的场景图。此外,我们构建了AeroEye-v1.0——一个新增五类交互类型的航拍视频数据集,突破现有数据集局限,为动态场景图生成提供全面基准。在ASPIRe和AeroEye-v1.0上的大量实验表明,所提THYME方法显著优于当前最优方法,在地面与航拍场景下均实现更优的场景理解性能。
原文摘要 · Abstract (English)
The rapid proliferation of video in applications such as autonomous driving, surveillance, and sports analytics necessitates robust methods for dynamic scene understanding. Despite advances in static scene graph generation and early attempts at video scene graph generation, previous methods often suffer from fragmented representations, failing to capture fine-grained spatial details and long-range temporal dependencies simultaneously. To address these limitations, we introduce the Temporal Hierarchical Cyclic Scene Graph (THYME) approach, which synergistically integrates hierarchical feature aggregation with cyclic temporal refinement to address these limitations. In particular, THYME effectively models multi-scale spatial context and enforces temporal consistency across frames, yielding more accurate and coherent scene graphs. In addition, we present AeroEye-v1.0, a novel aerial video dataset enriched with five types of interactivity that overcome the constraints of existing datasets and provide a comprehensive benchmark for dynamic scene graph generation. Empirically, extensive experiments on ASPIRe and AeroEye-v1.0 demonstrate that the proposed THYME approach outperforms state-of-the-art methods, offering improved scene understanding in ground-view and aerial scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。