DREAM通过双目标编码提升跨模态检索,精准匹配语言查询与视频内容。
DREAM: Extending Vision-Language Models with Dual-Objective Encoding for Cross-Modal Retrieval

- 采用混合语言建模,同时捕捉局部与全局语义。
- 在MSRVTT等数据集上达49.4%、49.7%、27.3%的R1新高。
- 适合关注视频检索与多模态对齐的研究者。
在媒体驱动的世界中,监控、教育和娱乐等领域视频内容呈指数增长,通过自然语言查询检索语义相关视频变得日益关键。早期系统依赖手工特征或浅层跨模态映射,难以捕捉复杂语义与时间动态。尽管大规模视觉-语言模型提升了跨模态对齐能力,但在建模细粒度时间依赖与细微语言结构方面仍存挑战。本文提出DREAM:双路径表示增强与对齐模型,通过强化视觉与文本编码解决上述问题。DREAM采用融合掩码与置换语言建模的目标,捕获局部与全局语言语义;视觉侧设计分层视觉编码器,结合级联分组注意力,通过多阶段标记交互与粗到精注意力优化整合空间与时间信息。在广泛使用的MSRVTT、MSVD和LSMDC基准数据集上全面验证,分别取得49.4%、49.7%和27.3%的R1新纪录。定性分析表明,模型能在帧间保持连贯注意力,并将复杂查询与动态视频内容精准对齐。结果凸显分层注意力与双目标文本建模在实现鲁棒、上下文感知视频检索中的有效性,为跨模态表示学习研究开辟新路径。
原文摘要 · Abstract (English)
In today's media-driven world, the exponential growth of video content across domains such as surveillance, education, and entertainment has made retrieving semantically relevant videos via natural language queries increasingly critical. Early video retrieval systems relied on handcrafted features or shallow cross-modal mappings, limiting their ability to capture complex semantics and temporal dynamics. While large-scale vision-language models have improved cross-modal alignment, challenges remain in modeling fine-grained temporal dependencies and nuanced linguistic structures. In this paper, we introduce DREAM: Dual-path Representation Enhancement and Alignment Model, a novel multimodal framework that addresses these limitations through enhanced visual and textual encoding. DREAM incorporates a hybrid language modeling strategy that combines masked and permuted language modeling objectives to capture both local and global linguistic semantics. On the visual side, we design a hierarchical vision encoder with cascaded group attention, which integrates spatial and temporal information through multi-stage token interaction and coarse-to-fine attention refinement. We validate DREAM through comprehensive evaluations on the widely-used MSRVTT, MSVD and LSMDC benchmark datasets, where it achieves new state-of-the-art R1 scores of 49.4%, 49.7% and 27.3%, respectively. Qualitative analyses further show the model's ability to maintain coherent attention across frames and align complex queries with dynamic video content. These findings underscore the effectiveness of hierarchical attention and dual-objective textual modeling in enabling robust, context-aware video retrieval, and pave the way for future research in advancing cross-modal representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。