arXiv:2503.23181cs.CV2025-03被引 2

改进视频定位推理策略,提升弱监督下的边界预测精度

Enhancing Weakly Supervised Video Grounding via Diverse Inference Strategies for Boundary and Prediction Selection

  • 用多高斯模型生成多样边界,替代固定偏移的单一边界
  • 引入质量评估机制,优化候选片段选择,提升定位准确率
  • 无需额外训练,在两个主流数据集上均取得显著提升

弱监督视频定位旨在未提供精确时间边界的情况下定位与查询相关的时序片段。现有方法主要依赖高斯分布生成候选片段,但忽略了推理阶段的边界预测和最优预测选择问题。在边界预测中,边界简单设定为均值两侧各半个标准差,可能无法捕捉最优边界;在最优预测选择中,过度依赖与其他候选片段的交集,未考虑各候选片段的质量差异。为此,本文探索多种推理策略,提出(1)从多个高斯分布中生成多样化边界的新型边界预测方法,(2)结合候选片段质量评估的新选择机制。在ActivityNet Captions和Charades-STA数据集上的大量实验验证了所提策略的有效性,性能提升无需额外训练。

原文摘要 · Abstract (English)

Weakly supervised video grounding aims to localize temporal boundaries relevant to a given query without explicit ground-truth temporal boundaries. While existing methods primarily use Gaussian-based proposals, they overlook the importance of (1) boundary prediction and (2) top-1 prediction selection during inference. In their boundary prediction, boundaries are simply set at half a standard deviation away from a Gaussian mean on both sides, which may not accurately capture the optimal boundaries. In the top-1 prediction process, these existing methods rely heavily on intersections with other proposals, without considering the varying quality of each proposal. To address these issues, we explore various inference strategies by introducing (1) novel boundary prediction methods to capture diverse boundaries from multiple Gaussians and (2) new selection methods that take proposal quality into account. Extensive experiments on the ActivityNet Captions and Charades-STA datasets validate the effectiveness of our inference strategies, demonstrating performance improvements without requiring additional training.

视频定位弱监督推理策略边界预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。