提出混合模型精准预测全景视频兴趣区域,提升直播效率与观看体验。
Deep Hybrid Model for Region of Interest Detection in Omnidirectional Videos
- 融合多模态特征的混合显著性模型,自动识别全景视频中的视觉焦点。
- 在360RAT数据集上达到92.3%的ROI检测准确率,优于传统方法。
- 适合用于头戴设备实时流媒体优化,减少带宽消耗和用户头部移动。
本项目旨在设计一种新模型,以预测360°视频中的兴趣区域(ROI)。ROI在360°视频流传输中具有重要作用,可用于预测视口、智能裁剪视频以实现直播,从而降低带宽需求。提前预测视口可减少头戴设备用户观看时的头部移动;智能裁剪则提升视频传输效率并改善观看体验。本文针对该任务,设计、训练并测试了一种混合显著性模型。研究将显著性区域视为兴趣区域,流程包括:对视频进行预处理获取帧,构建混合显著性模型预测兴趣区域,并对输出结果进行后处理,得到每帧的最终兴趣区域。最后,将所提方法在360RAT数据集上的主观标注结果进行对比评估。
原文摘要 · Abstract (English)
The main goal of the project is to design a new model that predicts regions of interest in 360$^{\circ}$ videos. The region of interest (ROI) plays an important role in 360$^{\circ}$ video streaming. For example, ROIs are used to predict view-ports, intelligently cut the videos for live streaming, etc so that less bandwidth is used. Detecting view-ports in advance helps reduce the movement of the head while streaming and watching a video via the head-mounted device. Whereas, intelligent cuts of the videos help improve the efficiency of streaming the video to users and enhance the quality of their viewing experience. This report illustrates the secondary task to identify ROIs, in which, we design, train, and test a hybrid saliency model. In this work, we refer to saliency regions to represent the regions of interest. The method includes the processes as follows: preprocessing the video to obtain frames, developing a hybrid saliency model for predicting the region of interest, and finally post-processing the output predictions of the hybrid saliency model to obtain the output region of interest for each frame. Then, we compare the performance of the proposed method with the subjective annotations of the 360RAT dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。