首次系统研究视频与文本的对齐能力,揭示模型性能关键影响因素。
Dynamic Reflections: Probing Video Representations with Text Alignment

- 通过测试时参数缩放规律,量化视频与文本对齐效果
- 发现强语义对齐与下游任务表现显著相关,尤其在通用理解上
- 构建时间推理挑战任务,为多模态模型提供新评估基准
跨模态表示对齐近年来被证明能揭示不同编码器在多种数据类型上的结构相似性与下游能力。尽管图像与文本对齐已取得进展,视频数据的时间特性在该领域仍鲜有探索。本文首次全面研究视频-文本表示对齐,探查现代视频与语言编码器的能力。结果表明:1)跨模态对齐高度依赖测试时输入数据的丰富性(静态图像对比多帧视频、单句描述对比多句集合),尤其在使用先进视频编码器时;我们提出参数化测试时缩放定律,对实证观察具有极强预测力。2)分析语义对齐与语义及非语义下游任务性能的相关性,初步表明与文本编码器强对齐可能关联于通用视频表征与理解能力。3)建立时间推理与跨模态对齐的关联,为视觉语言模型提供具有挑战性的测试平台。整体而言,本工作将视频-文本对齐引入为探测时空数据编码器表征能力的高效零样本工具。项目页面见 https://video-prh.github.io/
原文摘要 · Abstract (English)
The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types. While significant progress has been made in aligning images with text, the temporal nature of video data remains largely unexplored in this context. In this work, we conduct the first comprehensive study of video-text representation alignment, probing the capabilities of modern video and language encoders. Our findings reveal several key insights. First, we demonstrate that cross-modal alignment highly depends on the richness of both visual (static images vs. multi-frame videos) and text (single caption vs. a collection) data provided at test time, especially when using state-of-the-art video encoders. We propose parametric test-time scaling laws that capture this behavior and show remarkable predictive power against empirical observations. Secondly, we investigate the correlation between semantic alignment and performance on both semantic and non-semantic downstream tasks, providing initial evidence that strong alignment against text encoders may be linked to general-purpose video representation and understanding. Finally, we correlate temporal reasoning with cross-modal alignment providing a challenging test-bed for vision and language models. Overall, our work introduces video-text alignment as an informative zero-shot way to probe the representation power of different encoders for spatio-temporal data. Project page can be found at https://video-prh.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。