用双图增强多模态信息,提升视频段落描述生成效果
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
- 构建视频时序图与主题关联图,融合多模态信号和常识知识
- 在多个基准数据集上超越现有方法,显著提升生成质量
- 适合关注视频理解与生成、多模态融合的研究者
视频段落描述(VPC)旨在生成概括视频关键事件的段落式描述。尽管近期取得进展,仍面临有效利用视频中多模态信号及词汇长尾分布等挑战。本文提出一种新型多模态融合的标题生成框架,融合多种模态信息与外部知识库。框架构建两个图:一个'视频特定'时序图,捕捉主要事件及多模态信息与常识知识间的交互;另一个'主题图',表示特定主题下词语间的关联。这两个图作为共享编码器-解码器架构的Transformer网络输入。此外,引入节点选择模块,通过筛选图中相关节点提升解码效率。实验结果表明,该方法在多个基准数据集上表现优异。
原文摘要 · Abstract (English)
Video Paragraph Captioning (VPC) aims to generate paragraph captions that summarises key events within a video. Despite recent advancements, challenges persist, notably in effectively utilising multimodal signals inherent in videos and addressing the long-tail distribution of words. The paper introduces a novel multimodal integrated caption generation framework for VPC that leverages information from various modalities and external knowledge bases. Our framework constructs two graphs: a 'video-specific' temporal graph capturing major events and interactions between multimodal information and commonsense knowledge, and a 'theme graph' representing correlations between words of a specific theme. These graphs serve as input for a transformer network with a shared encoder-decoder architecture. We also introduce a node selection module to enhance decoding efficiency by selecting the most relevant nodes from the graphs. Our results demonstrate superior performance across benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。