针对视频场景图中长尾关系识别难问题,提出频率引导的多级推理模型。
Frequency-guided Multi-level Reasoning for Scene Graph Generation in Video

- 按关系频次设计专用分支,缓解梯度冲突
- 分离建模高频与低频关系,提升尾部类召回率
- 引入贝叶斯与高斯混合头,增强推理鲁棒性
视频场景图生成旨在为视频提供物体及其关系的结构化语义表示,以支持高层理解。然而,现有方法在处理长尾分布关系时仍存在局限。本文提出频率引导的关系多级推理(FReMuRe)模型,从机制层面增强对长尾关系的建模能力。通过引入关系特异性分支解决梯度冲突,实现更均衡且面向尾部关系的学习。设计频率感知的双分支谓词嵌入网络,分别建模高频与低频关系,并通过门控融合提升尾部类别召回率。同时提出两种可互换的关系分类头:贝叶斯头用于不确定性估计,新提出的高斯混合模型头增强类内多样性。实验表明,FReMuRe在Action Genome数据集上显著提升长尾关系的召回率与整体推理鲁棒性。
原文摘要 · Abstract (English)
Video Scene Graph Generation aims to obtain structured semantic representations of objects and their relationships in videos for high-level understanding. However, existing methods still have limitations in handling long-tail distributions. This paper proposes the Frequency-guided Relational Multi-level Reasoning (FReMuRe) model, which enhances the modeling ability of long-tail relationships from a mechanism perspective. We introduce relation-specific branches to deal gradient conflicts, yielding more balanced and tail-aware learning. And we design a frequency-aware dual-branch predicate embedding network to model high-frequency and low-frequency relationships separately and improve the recall rate of tail classes through gated fusion. Meanwhile, we propose two types of interchangeable relation classification heads: Bayesian Head for uncertainty estimation and new Gaussian Mixture Model Head to enhance intra-class diversity. Experimental results show that FReMuRe significantly improves the recall rate of long-tail relationships and overall reasoning robustness on the Action Genome dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。