arXiv:2603.08648cs.CV2026-03被引 4

让视频检索更连贯,通过预测状态变化提升长视频一致性

CAST: Modeling Visual State Transitions for Consistent Video Retrieval

论文配图:CAST: Modeling Visual State Transitions for Consistent Video Retrieval
图 1 · 摘自论文原文
  • 用视觉历史预测状态残差,显式建模视频状态演变
  • 在YouCook2和CrossTask上显著提升检索准确率
  • 适配主流模型,可为生成视频提供连贯性重排序

随着视频内容创作向长篇叙事发展,将短片段组合成连贯故事线愈发重要。然而,现有检索方法在推理时缺乏上下文感知,侧重局部语义匹配而忽视状态与身份一致性。为此,我们提出一致视频检索(CVR)任务,并构建涵盖YouCook2、COIN和CrossTask的诊断基准。提出CAST(上下文感知状态转移)模型,一种轻量级、即插即用的适配器,兼容多种冻结的视觉-语言嵌入空间。通过从视觉历史预测状态条件残差Δ,引入显式隐变量状态演化归纳偏置。大量实验表明,CAST在YouCook2和CrossTask上表现提升,于COIN保持竞争力,并在多种基础模型上持续优于零样本基线。此外,CAST可为黑箱视频生成候选结果(如Veo)提供有效重排序信号,促进更连贯的时序延续。

原文摘要 · Abstract (English)

As video content creation shifts toward long-form narratives, composing short clips into coherent storylines becomes increasingly important. However, prevailing retrieval formulations remain context-agnostic at inference time, prioritizing local semantic alignment while neglecting state and identity consistency. To address this structural limitation, we formalize the task of Consistent Video Retrieval (CVR) and introduce a diagnostic benchmark spanning YouCook2, COIN, and CrossTask. We propose CAST (Context-Aware State Transition), a lightweight, plug-and-play adapter compatible with diverse frozen vision-language embedding spaces. By predicting a state-conditioned residual update ($Δ$) from visual history, CAST introduces an explicit inductive bias for latent state evolution. Extensive experiments show that CAST improves performance on YouCook2 and CrossTask, remains competitive on COIN, and consistently outperforms zero-shot baselines across diverse foundation backbones. Furthermore, CAST provides a useful reranking signal for black-box video generation candidates (e.g., from Veo), promoting more temporally coherent continuations.

视频检索状态建模生成评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。