让视频检索理解隐含语义,通过动态原型捕捉概念关联
IMAGINE: Adaptive Schema-Imagery Enhanced Composition for Composed Video Retrieval

- 用动态多模态原型生成隐含语义(称作模式意象)
- 在三个基准上实现视频与图像检索的最新效果
- 适合需要理解语义联想的跨模态检索场景
组成视频检索(CVR)旨在找到与参考视频经修改文本调整后相匹配的目标视频。现有方法虽探索跨模态对应关系,但通常假设被修改对象在视频中明确出现。然而,修改文本常描述未显式呈现但通过语义相关视觉线索隐含表达的概念(如“蛋糕”暗示“生日派对”)。当前方法多依赖具体空间内的显式特征对齐,忽略了关键的潜在关联。为此,我们提出自适应模式-意象增强组合网络(IMAGINE)。不同于标准显式匹配,IMAGINE通过动态多模态原型具象化隐含语义(称为模式意象),这些原型捕捉共享潜在概念,自适应地调制视觉特征,有效将隐含引导注入检索过程。通过弥合显式视觉内容与隐含检索意图之间的差距,IMAGINE在三个广泛使用的基准上实现了视频检索(CVR)和组成图像检索(CIR)的最先进性能。
原文摘要 · Abstract (English)
Composed Video Retrieval (CVR) is designed to retrieve a target video that matches a reference video modified by a modification text. While existing methods explore cross-modal correspondences, they often assume modified objects appear directly in videos. However, modification texts frequently describe concepts not explicitly presented but implicitly expressed through semantically related visual cues (e.g., "cake" implying "birthday party"). Current approaches typically rely on aligning explicit feature representations within the concrete space, neglecting critical latent associations. To address this, we propose an adaptIve scheMa-ImAGery enhanced composItional NEtwork (IMAGINE). Unlike standard explicit matching, IMAGINE materializes implicit semantics (termed schema imagery) via dynamic multimodal prototypes. These prototypes capture shared latent concepts to adaptively modulate visual features, effectively injecting implicit guidance into the retrieval process. By bridging the gap between explicit visual contents and implicit retrieval intentions, IMAGINE achieves state-of-the-art performance in both CVR and Composed Image Retrieval (CIR) across three widely used benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。