模仿人类记忆层级,提升视频密集描述生成效果
HiCM$^2$: Hierarchical Compact Memory Modeling for Dense Video Captioning
- 构建分层紧凑记忆结构,模拟人类记忆回忆机制
- 在YouCook2和ViTT数据集上达到当前最优性能
- 适合关注视频理解与记忆建模的研究者
随着对现实世界视频挑战解决方案需求的增长,密集视频描述(DVC)研究日益受到重视。DVC旨在对未剪辑视频进行自动描述与定位。现有研究指出其挑战并引入基于先验知识的方法,如预训练与外部记忆。本文提出一种受人类层次化紧凑记忆启发的模型,通过构建分层记忆与分层读取模块,利用聚类记忆事件并结合大语言模型摘要,实现高效分层紧凑记忆。对比实验表明,该分层记忆召回机制显著提升DVC性能,在YouCook2和ViTT数据集上达到当前最优结果。
原文摘要 · Abstract (English)
With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the challenges of DVC and introduce improved methods utilizing prior knowledge, such as pre-training and external memory. In this research, we propose a model that leverages the prior knowledge of human-oriented hierarchical compact memory inspired by human memory hierarchy and cognition. To mimic human-like memory recall, we construct a hierarchical memory and a hierarchical memory reading module. We build an efficient hierarchical compact memory by employing clustering of memory events and summarization using large language models. Comparative experiments demonstrate that this hierarchical memory recall process improves the performance of DVC by achieving state-of-the-art performance on YouCook2 and ViTT datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。