提出多粒度对比协同生成模型,端到端解决长视频问答中的跨模态理解难题。
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
- 通过多粒度对比学习和单模态联合建模,提升视觉语义表征能力
- 将问答任务重构为生成式框架,在6个数据集上显著超越现有方法
- 适合需要精准跨模态推理的长视频理解场景
长期视频问答(VideoQA)是一项挑战性任务,要求对未剪辑的长视频进行语义理解,并回答多样化自由形式问题,同时强调全面的跨模态推理以获得精确答案。传统方法通常依赖现成特征提取器,虽降低计算开销,但导致领域无关、模态无关的表示;且单模态理解与跨模态交互之间的梯度阻塞,阻碍可靠答案生成。相比之下,新兴视频-语言预训练模型可实现低成本端到端建模,但在领域特定推理能力不足,且任务设定存在差异。为此,我们提出一种完全端到端的长期视频问答解决方案:多粒度对比协同生成(MCG)模型。为获取具备高视觉概念判别性的表征,我们在片段-骨架架构上引入联合单模态建模(JUM),并利用多粒度对比学习(MCL)挖掘内在或显式存在的语义对应关系。为缓解任务设定差异问题,提出跨模态协同生成(CCG)模块,将视频问答重构为生成任务而非传统分类方案,赋予模型跨模态高语义融合与生成能力,实现理性推理与答案生成。在六个公开可用的VideoQA数据集上的大量实验验证了所提方法的优势。
原文摘要 · Abstract (English)
Long-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-form questions, simultaneously emphasizing comprehensive cross-modal reasoning to yield precise answers. The canonical approaches often rely on off-the-shelf feature extractors to detour the expensive computation overhead, but often result in domain-independent modality-unrelated representations. Furthermore, the inherent gradient blocking between unimodal comprehension and cross-modal interaction hinders reliable answer generation. In contrast, recent emerging successful video-language pre-training models enable cost-effective end-to-end modeling but fall short in domain-specific ratiocination and exhibit disparities in task formulation. Toward this end, we present an entirely end-to-end solution for long-term VideoQA: Multi-granularity Contrastive cross-modal collaborative Generation (MCG) model. To derive discriminative representations possessing high visual concepts, we introduce Joint Unimodal Modeling (JUM) on a clip-bone architecture and leverage Multi-granularity Contrastive Learning (MCL) to harness the intrinsically or explicitly exhibited semantic correspondences. To alleviate the task formulation discrepancy problem, we propose a Cross-modal Collaborative Generation (CCG) module to reformulate VideoQA as a generative task instead of the conventional classification scheme, empowering the model with the capability for cross-modal high-semantic fusion and generation so as to rationalize and answer. Extensive experiments conducted on six publicly available VideoQA datasets underscore the superiority of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。