arXiv:2507.15504cs.CV2025-07ICCV被引 7

通过量化三种不确定性,让视频检索系统主动提问并逐步优化查询。

Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty Minimization

  • 用三个可计算指标量化文本模糊、映射不确定和帧质量差问题。
  • 在MSR-VTT-1k数据集上,10轮交互后召回率达69.2%。
  • 适合需要高精度视频检索的交互式应用,如智能搜索与内容推荐。

尽管近期取得进展,文本到视频检索(TVR)仍受多重固有不确定性制约,如文本查询模糊、文本-视频映射不明确以及视频帧质量低。尽管交互式系统通过提出澄清问题来改进用户意图,但现有方法多依赖启发式策略,未显式量化不确定性,限制了效果。为此,我们提出UMIVR框架,通过可解释、无需训练的度量方式,显式量化三类关键不确定性:基于语义熵的文本模糊度评分(TAS)、基于Jensen-Shannon散度的映射不确定性评分(MUS),以及基于时间质量的帧采样器(TQFS)。该框架根据这些度量自适应生成针对性澄清问题,迭代优化用户查询,显著降低检索歧义。在多个基准上的实验验证了其有效性,在MSR-VTT-1k数据集上,经10轮交互后,召回率@1达到69.2%,为交互式电视检索建立了不确定性最小化的新范式。

原文摘要 · Abstract (English)

Despite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and low-quality video frames. Although interactive systems have emerged to address these challenges by refining user intent through clarifying questions, current methods typically rely on heuristic or ad-hoc strategies without explicitly quantifying these uncertainties, limiting their effectiveness. Motivated by this gap, we propose UMIVR, an Uncertainty-Minimizing Interactive Text-to-Video Retrieval framework that explicitly quantifies three critical uncertainties-text ambiguity, mapping uncertainty, and frame uncertainty-via principled, training-free metrics: semantic entropy-based Text Ambiguity Score (TAS), Jensen-Shannon divergence-based Mapping Uncertainty Score (MUS), and a Temporal Quality-based Frame Sampler (TQFS). By adaptively generating targeted clarifying questions guided by these uncertainty measures, UMIVR iteratively refines user queries, significantly reducing retrieval ambiguity. Extensive experiments on multiple benchmarks validate UMIVR's effectiveness, achieving notable gains in Recall@1 (69.2\% after 10 interactive rounds) on the MSR-VTT-1k dataset, thereby establishing an uncertainty-minimizing foundation for interactive TVR.

视频检索交互式系统不确定性量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。