利用眼动追踪数据提升弱监督视频显著性目标检测效果
From Sight to Insight: Unleashing Eye-Tracking in Weakly Supervised Video Salient Object Detection
- 引入眼动注视点信息,通过位置与语义嵌入模块引导特征学习
- 在5个基准上超越现有方法,平均F-measure提升超过3%
- 适合关注人眼视觉规律与弱监督视频分析的研究者
眼动追踪视频显著性预测(VSP)和视频显著性目标检测(VSOD)均聚焦于视频中最具吸引力的物体,分别以预测热图和像素级显著性掩码形式呈现结果。实际应用中,眼动追踪标注更易获取,且与人眼真实视觉模式高度一致。本文旨在引入注视信息,在弱监督条件下辅助视频显著性目标检测。一方面,提出位置与语义嵌入(PSE)模块,在特征学习过程中提供位置与语义引导;另一方面,从特征选择与对比两个角度实现弱监督下的时空特征建模。设计了带语义与局部性约束的语义与局部性查询(SLQ)竞争器,有效选取最匹配的物体查询用于时空建模;同时提出一种内部-外部混合对比(IIMC)模型,通过视频内与跨视频对比学习增强弱监督下的时空建模能力。在五个主流VSOD基准上的实验结果表明,所提模型在多种评估指标上优于其他竞争方法。
原文摘要 · Abstract (English)
The eye-tracking video saliency prediction (VSP) task and video salient object detection (VSOD) task both focus on the most attractive objects in video and show the result in the form of predictive heatmaps and pixel-level saliency masks, respectively. In practical applications, eye tracker annotations are more readily obtainable and align closely with the authentic visual patterns of human eyes. Therefore, this paper aims to introduce fixation information to assist the detection of video salient objects under weak supervision. On the one hand, we ponder how to better explore and utilize the information provided by fixation, and then propose a Position and Semantic Embedding (PSE) module to provide location and semantic guidance during the feature learning process. On the other hand, we achieve spatiotemporal feature modeling under weak supervision from the aspects of feature selection and feature contrast. A Semantics and Locality Query (SLQ) Competitor with semantic and locality constraints is designed to effectively select the most matching and accurate object query for spatiotemporal modeling. In addition, an Intra-Inter Mixed Contrastive (IIMC) model improves the spatiotemporal modeling capabilities under weak supervision by forming an intra-video and inter-video contrastive learning paradigm. Experimental results on five popular VSOD benchmarks indicate that our model outperforms other competitors on various evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。