arXiv:2508.12084cs.CVcs.AI2025-08ICCV被引 3

用扩散模型生成多样且合理的视频事件边界,突破传统单一预测局限。

Generic Event Boundary Detection via Denoising Diffusion

  • 基于扩散模型,从时序自相似性编码中迭代生成事件边界。
  • 在Kinetics-GEBD和TAPOS上实现强性能,输出多样且合理的边界。
  • 引入新评估指标,兼顾结果多样性与准确性,适合视频理解研究者。

通用事件边界检测(GEBD)旨在识别视频中的自然事件边界,将其分割为独立且有意义的片段。尽管事件边界具有固有的主观性,但以往方法仅关注确定性预测,忽略了合理解的多样性。本文提出一种基于扩散的边界检测模型DiffGEBD,从生成视角解决GEBD问题。该模型通过时序自相似性编码相邻帧间的显著变化,并在条件特征引导下,逐步将随机噪声去噪为合理的事件边界。无分类器指引(classifier-free guidance)可控制生成结果的多样性。此外,我们引入新的评估指标,综合考量预测结果的多样性与保真度。实验表明,该方法在两个标准基准数据集Kinetics-GEBD和TAPOS上均表现优异,能生成多样且合理的事件边界。

原文摘要 · Abstract (English)

Generic event boundary detection (GEBD) aims to identify natural boundaries in a video, segmenting it into distinct and meaningful chunks. Despite the inherent subjectivity of event boundaries, previous methods have focused on deterministic predictions, overlooking the diversity of plausible solutions. In this paper, we introduce a novel diffusion-based boundary detection model, dubbed DiffGEBD, that tackles the problem of GEBD from a generative perspective. The proposed model encodes relevant changes across adjacent frames via temporal self-similarity and then iteratively decodes random noise into plausible event boundaries being conditioned on the encoded features. Classifier-free guidance allows the degree of diversity to be controlled in denoising diffusion. In addition, we introduce a new evaluation metric to assess the quality of predictions considering both diversity and fidelity. Experiments show that our method achieves strong performance on two standard benchmarks, Kinetics-GEBD and TAPOS, generating diverse and plausible event boundaries.

视频理解扩散模型事件检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。