区分视觉实体与场景的多模态特征,精准识别视频场景切换点。
Modality-Aware Shot Relating and Comparing for Video Scene Detection
- 分视角建模实体与场景信息,捕捉长短期镜头关联
- 通过相似性卷积明确编码前后镜头差异,定位场景结束帧
- 适用于多种视频数据集,提升场景检测精度
视频场景检测需判断每段镜头及其周围是否属于同一场景,关键在于精确关联多模态线索(如视觉实体与场景)并比较镜头间的语义变化。现有方法常平等地处理多模态语义,忽视镜头两侧上下文差异,导致性能受限。本文提出模态感知的镜头关联与比较方法(MASRC),依据视觉实体和场景模态的特性对镜头进行关联,并通过多镜头相似性比较显式编码场景变化。具体而言,从实体语义中挖掘长期镜头关联,从场景语义中揭示短期关联,从而学习具有场景内一致性与跨场景差异性的镜头特征。在此基础上,利用相似性卷积编码目标镜头前后镜头的关系,辅助识别场景结束帧。在多个公开基准数据集上的实验表明,MASRC显著提升视频场景检测性能。
原文摘要 · Abstract (English)
Video scene detection involves assessing whether each shot and its surroundings belong to the same scene. Achieving this requires meticulously correlating multi-modal cues, $\it{e.g.}$ visual entity and place modalities, among shots and comparing semantic changes around each shot. However, most methods treat multi-modal semantics equally and do not examine contextual differences between the two sides of a shot, leading to sub-optimal detection performance. In this paper, we propose the $\bf{M}$odality-$\bf{A}$ware $\bf{S}$hot $\bf{R}$elating and $\bf{C}$omparing approach (MASRC), which enables relating shots per their own characteristics of visual entity and place modalities, as well as comparing multi-shots similarities to have scene changes explicitly encoded. Specifically, to fully harness the potential of visual entity and place modalities in modeling shot relations, we mine long-term shot correlations from entity semantics while simultaneously revealing short-term shot correlations from place semantics. In this way, we can learn distinctive shot features that consolidate coherence within scenes and amplify distinguishability across scenes. Once equipped with distinctive shot features, we further encode the relations between preceding and succeeding shots of each target shot by similarity convolution, aiding in the identification of scene ending shots. We validate the broad applicability of the proposed components in MASRC. Extensive experimental results on public benchmark datasets demonstrate that the proposed MASRC significantly advances video scene detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。