通过运动引导语义对齐,提升动态场景图生成中细粒度关系建模能力。
MOSA: Motion-Guided Semantic Alignment for Dynamic Scene Graph Generation

- 用运动特征提取器捕捉距离、速度等对象对运动属性
- 融合运动与空间特征生成具备运动感知的关系表示
- 通过跨模态对齐和类别加权损失,改善长尾关系识别
动态场景图生成(DSGG)旨在对视频序列中的物体及其动态交互进行结构化建模,以实现高层语义理解。然而现有方法在细粒度关系建模、语义表征利用以及尾部关系建模方面存在不足。为此,本文提出一种运动引导语义对齐方法(MoSA)。首先,运动特征提取器(MFE)编码对象对的距离、速度、运动持续性及方向一致性等运动属性;随后,这些运动属性通过运动引导交互模块(MIM)与空间关系特征融合,生成运动感知的关系表示。为进一步增强语义区分能力,引入跨模态动作语义匹配(ASM)机制,对齐视觉关系特征与关系类别的文本嵌入。最后,采用类别加权损失策略,强化尾部关系的学习。大量严谨测试表明,MoSA在Action Genome数据集上表现最优。
原文摘要 · Abstract (English)
Dynamic Scene Graph Generation (DSGG) aims to structurally model objects and their dynamic interactions in video sequences for high-level semantic understanding. However, existing methods struggle with fine-grained relationship modeling, semantic representation utilization, and the ability to model tail relationships. To address these issues, this paper proposes a motion-guided semantic alignment method for DSGG (MoSA). First, a Motion Feature Extractor (MFE) encodes object-pair motion attributes such as distance, velocity, motion persistence, and directional consistency. Then, these motion attributes are fused with spatial relationship features through the Motion-guided Interaction Module (MIM) to generate motion-aware relationship representations. To further enhance semantic discrimination capabilities, the cross-modal Action Semantic Matching (ASM) mechanism aligns visual relationship features with text embeddings of relationship categories. Finally, a category-weighted loss strategy is introduced to emphasize learning of tail relationships. Extensive and rigorous testing shows that MoSA performs optimally on the Action Genome dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。