提出一到多对齐机制,让视频检索更精准
TokenBinder: Text-Video Retrieval with One-to-Many Alignment Paradigm

- 采用一到多对比对齐,模仿人类选物时的比较思维
- 在6个数据集上显著超越现有最优方法
- 适合需要精细匹配的视频搜索场景
文本-视频检索(TVR)通常通过粗粒度、细粒度或结合方式对齐文本与视频特征。然而,这些框架大多采用一(查询)对一(候选)的对齐范式,难以区分候选间的细微差异,导致频繁误匹配。受人类认知科学中比较判断的启发——决策通过直接比较而非独立评估做出,我们提出 TokenBinder。该创新的两阶段框架引入一种新颖的一对多粗细结合对齐范式,模拟人类从大量物品中识别特定项的认知过程。方法采用具有复杂交叉注意力机制的聚焦视图融合网络,动态对齐并比较多个视频的特征,以捕捉更细微的语义差异和上下文变化。在六个基准数据集上的大量实验表明,TokenBinder 显著优于现有最先进方法,验证了其在弥合模态内与模态间信息鸿沟方面的鲁棒性与有效性。
原文摘要 · Abstract (English)
Text-Video Retrieval (TVR) methods typically match query-candidate pairs by aligning text and video features in coarse-grained, fine-grained, or combined (coarse-to-fine) manners. However, these frameworks predominantly employ a one(query)-to-one(candidate) alignment paradigm, which struggles to discern nuanced differences among candidates, leading to frequent mismatches. Inspired by Comparative Judgement in human cognitive science, where decisions are made by directly comparing items rather than evaluating them independently, we propose TokenBinder. This innovative two-stage TVR framework introduces a novel one-to-many coarse-to-fine alignment paradigm, imitating the human cognitive process of identifying specific items within a large collection. Our method employs a Focused-view Fusion Network with a sophisticated cross-attention mechanism, dynamically aligning and comparing features across multiple videos to capture finer nuances and contextual variations. Extensive experiments on six benchmark datasets confirm that TokenBinder substantially outperforms existing state-of-the-art methods. These results demonstrate its robustness and the effectiveness of its fine-grained alignment in bridging intra- and inter-modality information gaps in TVR tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。