基于Transformer的全景视频显著性预测模型,精度显著提升。
SalFormer360: a transformer-based saliency estimation model for 360-degree videos
- 用SegFormer编码器+自定义解码器,结合视角中心偏置建模用户注意力
- 在三个主流数据集上显著超越现有方法,最高提升18.6%
- 适合做全景视频优化、视口预测等沉浸式应用的研究者参考
显著性估计近年来受到广泛关注,尤其在360度视频中对视口预测和沉浸式内容优化具有重要意义。本文提出SalFormer360,一种基于Transformer架构的360度视频显著性估计新模型。该模型结合了原有的SegFormer编码器并进行了微调以适应360度内容,同时设计了定制化解码器。为进一步提升预测精度,引入了观看中心偏置(Viewing Center Bias)机制,反映用户在360度环境中的注意力分布。在三个最大规模的显著性估计基准数据集上的大量实验表明,SalFormer360优于现有最先进方法:在Sport360上皮尔逊相关系数提升8.4%,PVS-HM上提升2.5%,VR-EyeTracking上提升18.6%。
原文摘要 · Abstract (English)
Saliency estimation has received growing attention in recent years due to its importance in a wide range of applications. In the context of 360-degree video, it has been particularly valuable for tasks such as viewport prediction and immersive content optimization. In this paper, we propose SalFormer360, a novel saliency estimation model for 360-degree videos built on a transformer-based architecture. Our approach is based on the combination of an existing encoder architecture, SegFormer, and a custom decoder. The SegFormer model was originally developed for 2D segmentation tasks, and it has been fine-tuned to adapt it to 360-degree content. To further enhance prediction accuracy in our model, we incorporated Viewing Center Bias to reflect user attention in 360-degree environments. Extensive experiments on the three largest benchmark datasets for saliency estimation demonstrate that SalFormer360 outperforms existing state-of-the-art methods. In terms of Pearson Correlation Coefficient, our model achieves 8.4% higher performance on Sport360, 2.5% on PVS-HM, and 18.6% on VR-EyeTracking compared to previous state-of-the-art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。