让视频语言模型更灵活,能应对任意文本输入。
Rethinking Video-Language Model from the Language Input Perspective

- 自动生成正负文本,挖掘细粒度语义
- 用视频引导跨模态对齐,提升性能
- 适配主流VLM,无需重训练
受大语言模型浪潮推动,视频-语言模型(VLMs)成为连接视频与文本的重要技术,但现有方法大多隐含假设:所有文本均预先定义于特定模板。然而在真实场景中,这种假设难以满足——1)预定义所有文本耗时耗力;2)模板限制严苛,用户体验差。研究发现,相同视频输入下,语义相似但模板不同的文本会导致性能差异。为此,本文提出一种即插即用框架,助力各类VLM方法实现更全面的视频-文本对齐。首先,从原始文本生成正负样本以聚焦特定文本成分;其次,设计基于属性的文本推理策略,挖掘生成文本的细粒度语义;最后,利用视频作为引导,通过自加权损失实现跨模态桥接。大量实验证明,该方法可作为通用模块显著提升当前先进VLM的性能。
原文摘要 · Abstract (English)
Driven by the wave of large language models, Video-Language Models (VLMs) have become a significant yet challenging technology to bridge the gap between videos and texts. Although previous VLM works have made significant progress, almost all of them implicitly assume that all the texts are predefined by the specific template. In real-world applications, such a strict assumption is impossible to satisfy since 1) predefining all the texts is extremely time-consuming and labor-intensive. 2) these predefined text inputs are too restrictive and user-unfriendly, limiting their applications. It is observed that given a video input, texts with similar semantics but different templates lead to various performances. To this end, in this paper, we propose a novel plug-and-play framework for various VLM-based methods to fully bridge videos and texts. Specifically, we first generate positive and negative texts from the original ones to target specific text components. Then, we propose an attribute-based text reasoning strategy to mine fine-grained textual semantics of generated texts. Finally, we utilize videos as guidance to conduct cross-modal bridging by designing a self-weighted loss. Extensive experiments show that the proposed method can serve as the plug-and-play module to effectively improve the performance of state-of-the-art VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。