首个面向第一人称空间视频的感知质量评估数据集与模型。
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
- 构建首个第一人称空间视频质量数据库,含600段苹果Vision Pro拍摄视频。
- 提出ESVQAnet模型,在新任务上超越16个主流视频质量评估模型。
- 适用于沉浸式XR体验优化,适合多媒体质量评估研究者使用。
随着扩展现实(XR)的快速发展,第一人称空间拍摄与显示技术进一步提升了用户的沉浸感与参与度,带来了更吸引人、更具交互性的体验。评估第一人称空间视频的体验质量(QoE)对保障高质量观看至关重要,但相关研究仍较匮乏。本文引入具身体验概念,研究第一人称空间视频的具身感知质量评估新问题。具体地,我们构建了首个第一人称空间视频质量评估数据库(ESVQAD),包含600段使用Apple Vision Pro拍摄的第一人称空间视频及其对应的平均意见得分(MOS)。此外,我们提出一种新型多维双目特征融合模型ESVQAnet,整合双目空间、运动与语义特征以预测整体感知质量。实验结果表明,ESVQAnet在具身感知质量评估任务上显著优于16个主流视频质量评估模型,并在传统VQA任务上展现出强泛化能力。数据集与代码已开源:https://github.com/iamazxl/ESVQA。
原文摘要 · Abstract (English)
With the rapid development of eXtended Reality (XR), egocentric spatial shooting and display technologies have further enhanced immersion and engagement for users, delivering more captivating and interactive experiences. Assessing the quality of experience (QoE) of egocentric spatial videos is crucial to ensure a high-quality viewing experience. However, the corresponding research is still lacking. In this paper, we use the concept of embodied experience to highlight this more immersive experience and study the new problem, i.e., embodied perceptual quality assessment for egocentric spatial videos. Specifically, we introduce the first Egocentric Spatial Video Quality Assessment Database (ESVQAD), which comprises 600 egocentric spatial videos captured using the Apple Vision Pro and their corresponding mean opinion scores (MOSs). Furthermore, we propose a novel multi-dimensional binocular feature fusion model, termed ESVQAnet, which integrates binocular spatial, motion, and semantic features to predict the overall perceptual quality. Experimental results demonstrate the ESVQAnet significantly outperforms 16 state-of-the-art VQA models on the embodied perceptual quality assessment task, and exhibits strong generalization capability on traditional VQA tasks. The database and code are available at https://github.com/iamazxl/ESVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。