给视频时间定位结果加可靠区间,确保真答案在其中的概率至少为1-α。
Conformal Coverage Guarantees for Any Video Temporal Grounder

- 用事后校准的非符合度分数,将任意定位器转为带置信区间的输出。
- 在三个基准上实测覆盖率达目标值,点指标无法揭示的问题被暴露。
- 无需重训练或模型白盒,适用于黑箱视频语言模型等场景。
连续视频中的事件边界具有模糊性:同一查询-视频对多次标注,不同标注者标记的时间段重叠率常低于一半。因此,视频时间定位的真实标签应为区间分布,但现有定位器仅输出单一区间且不提供可靠性声明,导致部署时错误区间与正确区间无法区分。COVER提出一种事后、模型无关的封装方法,将任意定位器(包括训练好的局部定位器或黑箱视频-语言模型)转换为输出包含真实时刻的概率至少为1−α的时序区域。该方法通过在保留标签上校准时间非符合度分数的分位数,并据此扩展基础预测范围实现。保证为有限样本且分布自由,在可交换性假设下成立,无需重新训练或白盒访问。本文提出两种得分形式:适用于输出区间的双边边界扩展得分,以及适用于输出相关性信号的超水平集得分;并建立了针对定位任务的理论,量化了认证区域大小,分析了在事件长度条件下的覆盖率保持情况,以及当来自同一视频的时刻破坏可交换性时的退化现象。在三个基准和五个定位器上,实际覆盖率稳定接近目标值,校准过程揭示了点指标所掩盖的问题。
原文摘要 · Abstract (English)
Event boundaries in continuous video are ambiguous: re-annotate the same query-video pair and independent annotators mark moments that overlap by less than half on a large fraction of samples. The ground truth for video temporal grounding is therefore a distribution over intervals, yet every grounder returns a single interval with no statement of reliability, so at deployment a wrong interval is indistinguishable from a right one. COVER changes the output object: a post-hoc, model-agnostic wrapper that turns any grounder, a trained localizer or a black-box video--language model, into one that emits a temporal region containing the true moment with probability at least $1-α$, by calibrating the quantile of a temporal nonconformity score on held-out labels and widening the base prediction by that amount. The guarantee is finite-sample and distribution-free under exchangeability, and requires neither retraining nor white-box access. We give two score families, a two-sided boundary-widening score for grounders that emit an interval and a super-level-set score for grounders that emit a relevance signal, and develop theory specific to grounding that bounds how large the certified region becomes, when coverage survives conditioning on event length, and how it degrades when moments from one video break exchangeability. Across three benchmarks and five grounders, realized coverage tracks the target, and calibration exposes what point metrics hide.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。