arXiv:2503.15846cs.CV2025-03

用大模型重做动态场景图生成,提升准确性和实用性。

A Closer Look at Dynamic Scene Graph Generation In the Era of Multimodal Large Language Models

  • 改用自上而下推理定位,替代传统检测流程。
  • 提出时间关系集合预测,效率与性能双提升。
  • 引入重要性感知微调,生成更相关多样的关系。

动态场景图生成(DSGG)旨在捕捉视频中物体及其随时间演化的关系。尽管多模态大模型(MLLMs)快速发展,现有生成结果在实用性和质量上仍显不足。本文从任务设定与模型设计两方面重新审视DSGG:任务层面,指出当前以召回率为导向的评估存在严重精确率-召回率权衡及冗余关系生成问题,提出五个新指标以更全面评估生成质量;模型层面,探索直接使用MLLM进行场景图生成,通过三项改进建立强基线:(1)以自上而下‘先推理后定位’策略替代传统自下而上的流水线;(2)将帧级动态图重构为时间关系集合(TRS)预测,提升效率与性能;(3)引入重要性感知微调(IAF),促进更相关且多样化的关系生成。在Action Genome、VidVRD和PVSG数据集上的实验表明,该方法持续达到最优表现。

原文摘要 · Abstract (English)

Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graphs remain limited compared to the rapid advances in Multimodal Large Language Models (MLLMs). In this work, we revisit DSGG from two fundamental perspectives: task setup and model design. From the task setup perspective, we identify two key limitations of the current recall-oriented evaluation protocol: (i) a severe precision-recall trade-off, and (ii) uninformative and redundant relation generation. To better assess practical usefulness, we introduce five additional metrics that provide a more comprehensive evaluation of the quality of generated dynamic scene graphs. From the model design perspective, we explore directly using MLLMs for scene graph generation and establish a strong MLLM-based DSGG baseline through three design changes. First, we replace the conventional bottom-up pipeline with a top-down reason-then-locate strategy. Second, we reformulate frame-wise dynamic graphs as Temporal Relation Set (TRS) prediction, improving both efficiency and performance. Third, we introduce Importance-Aware Finetuning (IAF) to encourage more relevant and diverse relation generation. Extensive experiments on Action Genome, VidVRD, and PVSG show that our approach consistently achieves state-of-the-art performance.

动态场景图多模态大模型关系预测视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。