构建首个卫星视频时空全景场景图生成基准,助力动态地理场景理解
T-STAR: A Large-Scale Benchmark for Spatio-Temporal Panoptic Scene Graph Generation in Satellite Video

- 提出时空全景场景图生成任务,统一建模对象身份与时空关系
- 构建包含110万实例掩码、380万时空三元组的大规模数据集
- 适合研究遥感视频理解、时空关系建模的学者使用
结构化理解卫星视频对推动动态地理场景分析从低层感知到高层认知至关重要。为突破以目标为中心的感知局限,本文提出卫星视频中的时空全景场景图生成(TPSG)作为新基准任务。TPSG旨在生成由带有明确时间跨度的三元组<主体, 关系, 客体>组成的结构化图,通过联合建模一致的身份实例掩码与全景场景元素间的时空关系,描述动态地理场景。然而,当前尚无专用于卫星视频TPSG的数据集。且卫星视频中物体通常小而纹理弱,跨帧关联易受遮挡和背景杂乱干扰,关系语义高度耦合空间结构与时间演化,导致自然视频的TPSG模型难以直接应用。本文提出T-STAR,一个大规模基准数据集,涵盖39种细粒度物体类别和70种细粒度关系类别,包含超过110万实例掩码和超过380万时空三元组。为支持卫星视频中的TPSG,我们设计统一框架以增强跨帧实例一致性与时空关系预测。大量实验验证了T-STAR的重要性及所提框架的有效性,为未来结构化卫星视频理解研究建立了强基准。数据集与代码已公开于https://github.com/linlin-dev/T-STAR。
原文摘要 · Abstract (English)
Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from low-level perception to high-level cognition. To move beyond object-centric perception, this paper introduces spatio-temporal panoptic scene graph generation (TPSG) in satellite video as a new benchmark task. TPSG aims to generate a structured graph composed of a set of triplets <subject, relationship, object> with explicit temporal spans, thereby describing dynamic geospatial scenes by jointly modeling identity-consistent instance masks and spatio-temporal relationships among panoptic scene elements. However, there is still no dedicated dataset for TPSG in satellite video. Moreover, TPSG in satellite video is intrinsically challenging, as objects are often small and weakly textured, cross-frame association is easily disrupted by occlusion and background clutter, and relationship semantics are highly coupled with spatial structure and temporal evolution. Consequently, TPSG models developed for natural videos are not directly applicable to satellite video. This paper presents T-STAR, a large-scale benchmark dataset for TPSG in satellite video, comprising over 1.1 million instance masks and over 3.8 million spatio-temporal triplets across 39 fine-grained object categories and 70 fine-grained relationship categories. To enable TPSG in satellite video, we propose a unified framework to enhance cross-frame instance consistency and spatio-temporal relationship prediction. Extensive experiments demonstrate the significance of T-STAR and the effectiveness of the proposed framework, establishing a strong benchmark for future research on structured satellite video understanding. The dataset and code are available at https://github.com/linlin-dev/T-STAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。