arXiv:2503.00548cs.CVcs.MM2025-03CVPR被引 7

提出双路去偏框架,提升视频场景图生成的准确性与公平性。

Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing

  • 通过时序记忆增强整合视觉特征,缓解视觉偏差。
  • 利用三元组语义信息迭代融合,显著降低语义偏差。
  • 在半约束条件下,关键指标提升13.1%,适合动态场景分析任务。

视频场景图生成(VidSGG)旨在通过逐帧分析并融合视觉与语义信息,捕捉实体间的动态关系。然而,现有方法受显著偏差影响,导致预测失准。为此,我们提出一种视觉与语义感知(VISA)框架,实现无偏视频场景图生成。VISA通过增强时序记忆的视觉特征整合,缓解视觉偏差;同时,通过将物体特征与三元组关系推导的全面语义信息进行迭代融合,有效减少语义偏差。该双路去偏策略使复杂场景动态的表征更趋公正。大量实验表明,该方法在多个基准上表现优异,在半约束条件下的SGCLS任务中,mR@20和mR@50指标分别提升13.1%以上,显著优于现有无偏方法。

原文摘要 · Abstract (English)

Video Scene Graph Generation (VidSGG) aims to capture dynamic relationships among entities by sequentially analyzing video frames and integrating visual and semantic information. However, VidSGG is challenged by significant biases that skew predictions. To mitigate these biases, we propose a VIsual and Semantic Awareness (VISA) framework for unbiased VidSGG. VISA addresses visual bias through memory-enhanced temporal integration that enhances object representations and concurrently reduces semantic bias by iteratively integrating object features with comprehensive semantic information derived from triplet relationships. This visual-semantics dual debiasing approach results in more unbiased representations of complex scene dynamics. Extensive experiments demonstrate the effectiveness of our method, where VISA outperforms existing unbiased VidSGG approaches by a substantial margin (e.g., +13.1% improvement in mR@20 and mR@50 for the SGCLS task under Semi Constraint).

视频生成场景图去偏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。