arXiv:2412.06182cs.CV2024-12被引 17

将长视频转化为分层文本,提升理解精度与效率

Towards Long Video Understanding via Fine-detailed Video Story Generation

  • 自底向上逐级解析视频片段,实现细粒度时序建模
  • 双层级去冗余机制,有效降低视觉与文本干扰
  • 无需微调即可适配多种任务,适用性广

长视频理解已成为计算机视觉中的关键任务,推动了从监控到内容检索等众多应用的发展。现有方法在处理长视频时面临两大挑战:复杂的长时程关系建模和冗余信息干扰。为此,我们提出细粒度视频故事生成(FDVS),将长视频转化为详细的文本表征。具体而言,为实现对长时序内容的细粒度建模,我们设计了一种自底向上的视频解读机制,逐步从片段到整体解析视频内容。为避免视频中冗余信息的干扰,引入语义冗余消除机制,在视觉与文本层面同时去除冗余。该方法生成包含多粒度信息的分层文本表征,使FDVS可在不进行微调的情况下应用于多种任务。我们在八个涵盖三个任务的数据集上进行了评估,结果验证了方法的有效性与通用性。

原文摘要 · Abstract (English)

Long video understanding has become a critical task in computer vision, driving advancements across numerous applications from surveillance to content retrieval. Existing video understanding methods suffer from two challenges when dealing with long video understanding: intricate long-context relationship modeling and interference from redundancy. To tackle these challenges, we introduce Fine-Detailed Video Story generation (FDVS), which interprets long videos into detailed textual representations. Specifically, to achieve fine-grained modeling of long-temporal content, we propose a Bottom-up Video Interpretation Mechanism that progressively interprets video content from clips to video. To avoid interference from redundant information in videos, we introduce a Semantic Redundancy Reduction mechanism that removes redundancy at both the visual and textual levels. Our method transforms long videos into hierarchical textual representations that contain multi-granularity information of the video. With these representations, FDVS is applicable to various tasks without any fine-tuning. We evaluate the proposed method across eight datasets spanning three tasks. The performance demonstrates the effectiveness and versatility of our method.

视频理解长视频文本生成去冗余

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。