用生成对抗网络和球面卷积,提升全景视频显著性检测精度。
Sphere-GAN: a GAN-based Approach for Saliency Estimation in 360° Videos
- 基于球面卷积的GAN架构,适配360°视频的全景特性。
- 在公开数据集上显著优于现有最先进模型,提升显著性预测准确率。
- 适合做全景视频压缩、渲染优化的研究者或工程师参考。
沉浸式应用的兴起推动了对360°图像与视频处理的新方法研究,其中显著性估计可识别视觉关注区域,进而优化处理流程。尽管2D内容的显著性估计已广泛研究,针对360°视频的方法仍很少。为此,本文提出Sphere-GAN,一种基于生成对抗网络(GAN)并结合球面卷积的360°视频显著性检测模型。通过在公开的360°视频显著性数据集上进行大量实验,结果表明Sphere-GAN在预测显著性图方面显著优于现有最先进模型。
原文摘要 · Abstract (English)
The recent success of immersive applications is pushing the research community to define new approaches to process 360° images and videos and optimize their transmission. Among these, saliency estimation provides a powerful tool that can be used to identify visually relevant areas and, consequently, adapt processing algorithms. Although saliency estimation has been widely investigated for 2D content, very few algorithms have been proposed for 360° saliency estimation. Towards this goal, we introduce Sphere-GAN, a saliency detection model for 360° videos that leverages a Generative Adversarial Network with spherical convolutions. Extensive experiments were conducted using a public 360° video saliency dataset, and the results demonstrate that Sphere-GAN outperforms state-of-the-art models in accurately predicting saliency maps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。