提出新训练框架,让视频场景图生成更公平、更抗偏差。
Towards Unbiased and Robust Spatio-Temporal Scene Graph Generation and Anticipation
- 用损失掩码和课程学习减少头部关系的主导影响
- 在Action Genome上显著提升尾部关系生成准确率
- 适合关注模型公平性与鲁棒性的视觉理解研究者
时空场景图(STSG)通过建模物体及其随时间演化的关系,为动态场景提供简洁而丰富的表示。然而,现实视觉关系常呈长尾分布,导致现有视频场景图生成(VidSGG)与场景图预测(SGA)方法产生偏差。为此,我们提出ImparTail,一种新型训练框架,利用损失掩码和课程学习缓解时空场景图生成与预测中的偏差。不同于以往通过增加网络结构学习无偏估计器的方法,我们设计了公正的训练目标,在学习过程中降低头部类别的主导作用,聚焦于低频尾部关系。基于课程的掩码生成策略使模型能随时间自适应调整偏差缓解策略,实现更均衡、更鲁棒的估计。为进一步评估模型在分布偏移下的性能,我们引入两个新任务:鲁棒时空场景图生成与鲁棒场景图预测,构成具有挑战性的基准测试。在Action Genome数据集上的大量实验表明,本方法在无偏性能与鲁棒性方面均优于现有基线。
原文摘要 · Abstract (English)
Spatio-Temporal Scene Graphs (STSGs) provide a concise and expressive representation of dynamic scenes by modeling objects and their evolving relationships over time. However, real-world visual relationships often exhibit a long-tailed distribution, causing existing methods for tasks like Video Scene Graph Generation (VidSGG) and Scene Graph Anticipation (SGA) to produce biased scene graphs. To this end, we propose ImparTail, a novel training framework that leverages loss masking and curriculum learning to mitigate bias in the generation and anticipation of spatio-temporal scene graphs. Unlike prior methods that add extra architectural components to learn unbiased estimators, we propose an impartial training objective that reduces the dominance of head classes during learning and focuses on underrepresented tail relationships. Our curriculum-driven mask generation strategy further empowers the model to adaptively adjust its bias mitigation strategy over time, enabling more balanced and robust estimations. To thoroughly assess performance under various distribution shifts, we also introduce two new tasks Robust Spatio-Temporal Scene Graph Generation and Robust Scene Graph Anticipation offering a challenging benchmark for evaluating the resilience of STSG models. Extensive experiments on the Action Genome dataset demonstrate the superior unbiased performance and robustness of our method compared to existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。