arXiv:2501.11653cs.CVcs.LG2025-01被引 2

用视觉语言模型理解复杂动态场景,无需大量训练参数。

Dynamic Scene Understanding from Vision-Language Representations

  • 将动态场景理解任务转化为结构化文本预测或特征拼接。
  • 在多个任务上达到当前最优性能,仅需极少可训练参数。
  • 适合需要轻量级、通用性场景理解的AI应用开发人员。

描绘复杂动态场景的图像自动解析极具挑战性,需兼具整体情境的高层理解与参与实体及其交互的细粒度识别。现有方法针对情景识别、人-人及人-物交互检测等子任务采用不同策略。然而,近期图像理解进展常借助大规模网络级视觉-语言(V&L)表示,避免特定任务工程设计。本文提出一种利用现代冻结式V&L表示进行动态场景理解的框架。通过将这些任务统一建模为预测和解析结构化文本,或直接拼接表示至现有模型输入,我们在保持极低可训练参数量的同时实现了领先性能。此外,对这些表示中动态知识的分析表明,更强大的新模型有效编码了动态场景语义,使该方法成为可能。

原文摘要 · Abstract (English)

Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their interactions. Current approaches use distinct methods tailored to sub-tasks such as Situation Recognition and detection of Human-Human and Human-Object Interactions. However, recent advances in image understanding have often leveraged web-scale vision-language (V&L) representations to obviate task-specific engineering. In this work, we propose a framework for dynamic scene understanding tasks by leveraging knowledge from modern, frozen V&L representations. By framing these tasks in a generic manner - as predicting and parsing structured text, or by directly concatenating representations to the input of existing models - we achieve state-of-the-art results while using a minimal number of trainable parameters relative to existing approaches. Moreover, our analysis of dynamic knowledge of these representations shows that recent, more powerful representations effectively encode dynamic scene semantics, making this approach newly possible.

视觉语言模型场景理解动态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。