提出首个360度视频文本驱动显著性检测方法与数据集。
TSalV360: A Method and Dataset for Text-driven Saliency Detection in 360-Degrees Videos
- 用视觉语言模型融合文本与视频,实现跨模态显著性定位。
- 在1.6万组数据上验证,显著性预测精度优于现有方法。
- 适合需要定制化内容推荐的全景视频应用开发人员。
本文针对360度视频中的文本驱动显著性检测任务,提出了TSV360数据集,包含16,000个等距矩形投影(ERP)帧、对应显著对象/事件的文本描述及其真实显著图。在此基础上,扩展并适配一种先进的基于视觉的方法,提出TSalV360方法,可利用用户提供的目标对象或事件的文本描述进行引导。该方法采用先进的视觉-语言模型进行多模态表征,结合相似性估计模块和视口时空交叉注意力机制,挖掘不同模态间的依赖关系。通过定量与定性评估,TSalV360在TSV360数据集上的表现优于当前最先进的视觉基方法,证明了其在360度视频中实现个性化文本驱动显著性检测的能力。
原文摘要 · Abstract (English)
In this paper, we deal with the task of text-driven saliency detection in 360-degrees videos. For this, we introduce the TSV360 dataset which includes 16,000 triplets of ERP frames, textual descriptions of salient objects/events in these frames, and the associated ground-truth saliency maps. Following, we extend and adapt a SOTA visual-based approach for 360-degrees video saliency detection, and develop the TSalV360 method that takes into account a user-provided text description of the desired objects and/or events. This method leverages a SOTA vision-language model for data representation and integrates a similarity estimation module and a viewport spatio-temporal cross-attention mechanism, to discover dependencies between the different data modalities. Quantitative and qualitative evaluations using the TSV360 dataset, showed the competitiveness of TSalV360 compared to a SOTA visual-based approach and documented its competency to perform customized text-driven saliency detection in 360-degrees videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。