构建时序一致的动态场景图,端到端生成动作轨迹序列。
Temporally Consistent Dynamic Scene Graphs: An End-to-End Approach for Action Tracklet Generation
- 通过双向匹配与自适应解码器实现跨帧关系追踪
- 在三个数据集上时序召回率提升超60%
- 适合视频理解、自动驾驶等需要连续动作分析的场景
理解视频内容对活动识别、自主系统和人机交互等实际应用至关重要。虽然场景图能捕捉单帧中的空间关系,但将其扩展到跨视频序列的动态交互仍具挑战。为此,我们提出TCDSG——时序一致的动态场景图,一种端到端框架,可检测、追踪并链接跨时间的主谓宾关系,生成动作轨迹。该方法采用新颖的二分匹配机制,结合自适应解码器查询与反馈环路,确保长时间序列的时序一致性与鲁棒性。实验表明,在Action Genome、OpenPVSG和MEVA数据集上,时序召回率@k提升超过60%;同时首次为MEVA数据集增加持久对象ID标注,支持完整轨迹生成。本工作实现了空间与时间动态的无缝融合,推动多帧视频分析新标准,为监控、自主导航等高影响力应用开辟新路径。
原文摘要 · Abstract (English)
Understanding video content is pivotal for advancing real-world applications like activity recognition, autonomous systems, and human-computer interaction. While scene graphs are adept at capturing spatial relationships between objects in individual frames, extending these representations to capture dynamic interactions across video sequences remains a significant challenge. To address this, we present TCDSG, Temporally Consistent Dynamic Scene Graphs, an innovative end-to-end framework that detects, tracks, and links subject-object relationships across time, generating action tracklets, temporally consistent sequences of entities and their interactions. Our approach leverages a novel bipartite matching mechanism, enhanced by adaptive decoder queries and feedback loops, ensuring temporal coherence and robust tracking over extended sequences. This method not only establishes a new benchmark by achieving over 60% improvement in temporal recall@k on the Action Genome, OpenPVSG, and MEVA datasets but also pioneers the augmentation of MEVA with persistent object ID annotations for comprehensive tracklet generation. By seamlessly integrating spatial and temporal dynamics, our work sets a new standard in multi-frame video analysis, opening new avenues for high-impact applications in surveillance, autonomous navigation, and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。