用查询分解提升跨语言图文检索,零样本下更准找视频。
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
- 用大模型知识自动拆解复杂查询,挖掘隐藏事件信息。
- 在两个数据集上超越多个顶尖基线,音频融合显著提效。
- 支持多语言、多模态输入,适合跨域视频检索研究者。
近期方法已展示出从大型语言模型(LLMs)和视觉-语言模型(VLMs)中提取并利用参数化知识的强大能力。本文提出一种用于零样本多语言文本到视频检索的查询到事件分解方法(Q2E),可适配不同数据集、领域、LLM或VLM。我们的方法通过利用嵌入在LLMs和VLMs中的知识,对原本过于简化的用户查询进行分解,从而提升对真实世界复杂事件的理解与检索能力。我们还展示了该方法在视觉与语音输入上的适用性。为融合多模态知识,采用基于熵的融合评分策略实现零样本融合。在两个多样化数据集和多种评估指标下,Q2E表现优于多个先进基线。实验还表明,引入音频信息能显著提升文本到视频检索效果。代码与数据已公开,供后续研究使用。
原文摘要 · Abstract (English)
Recent approaches have shown impressive proficiency in extracting and leveraging parametric knowledge from Large-Language Models (LLMs) and Vision-Language Models (VLMs). In this work, we consider how we can improve the identification and retrieval of videos related to complex real-world events by automatically extracting latent parametric knowledge about those events. We present Q2E: a Query-to-Event decomposition method for zero-shot multilingual text-to-video retrieval, adaptable across datasets, domains, LLMs, or VLMs. Our approach demonstrates that we can enhance the understanding of otherwise overly simplified human queries by decomposing the query using the knowledge embedded in LLMs and VLMs. We additionally show how to apply our approach to both visual and speech-based inputs. To combine this varied multimodal knowledge, we adopt entropy-based fusion scoring for zero-shot fusion. Through evaluations on two diverse datasets and multiple retrieval metrics, we demonstrate that Q2E outperforms several state-of-the-art baselines. Our evaluation also shows that integrating audio information can significantly improve text-to-video retrieval. We have released code and data for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。