用分布对齐提升文本到视频检索的不确定性感知能力
Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

- 将文本与视频嵌入建模为高斯分布,显式捕捉模态不确定性
- 通过截断迭代优化,使文本分布逼近目标视频分布,提升检索精度
- 支持校准的不确定性排名,适合需要可解释性的检索场景
本文提出分布对齐桥(DAB),将文本到视频检索重新定义为分布对齐任务,而非传统的确定性点匹配。通过将文本和视频嵌入建模为由均值和方差定义的高斯分布,DAB 显式建模了模态特有的不确定性。采用一种确定性的、受扩散启发的桥接机制,通过截断精炼过程迭代优化文本分布,使其逐步逼近目标视频分布。该方法将概率嵌入与分布变换统一为端到端可训练系统。为优化跨模态相似性,引入基于相对熵(Kullback-Leibler divergence)的分布感知对比损失。在 MSR-VTT、MSVD 与 VATEX 基准上的大量实验表明,DAB 显著优于现有概率与扩散基线方法,并通过桥接诱导的分布边界实现校准的不确定性感知排序。
原文摘要 · Abstract (English)
This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task rather than traditional deterministic point matching. By modeling both text and video embeddings as Gaussian distributions defined by mean and variance, DAB explicitly accounts for modality-specific uncertainty. We employ a deterministic, diffusion-inspired bridge to iteratively refine text distributions toward their target video distributions through a truncated refinement process. This approach unifies probabilistic embedding and distributional transformation into a cohesive, end-to-end trainable system. To optimize cross-modal similarity, we introduce a distribution-aware contrastive loss based on Kullback-Leibler divergence. Extensive evaluations on MSR-VTT, MSVD, and VATEX benchmarks confirm that DAB significantly outperforms existing probabilistic and diffusion-based baselines, while providing calibrated uncertainty-aware ranking through bridge-induced distributional margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。