arXiv:2506.10331cs.CVeess.IV2025-06中稿 · ICME 2025

构建首个UGC全景视频音画质量数据集并提出高效评估模型

Research on Audio-Visual Quality Assessment Dataset and Method for User-Generated Omnidirectional Video

  • 用5人拍摄2类全景相机,采集300段覆盖10场景的UGC全景音视频
  • 通过主观实验获取均值评分(MOS),验证模型在该数据集上表现最优
  • 提出音视频特征提取与融合框架,适用于社交平台等真实场景

针对元宇宙兴起背景下全景视频(ODVs)从专业生成内容(PGC)向用户生成内容(UGC)转变的趋势,当前音画质量评估(AVQA)研究仍显不足。为此,本文构建了首个基于UGC的全景音视频数据集:由5名参与者使用2种不同类型的全景相机拍摄300段视频,涵盖10类场景。通过主观实验获取各音视频序列的平均意见得分(MOS)。为进一步推动该领域发展,基于该数据集构建了一个有效的AVQA基准模型,包含视频特征提取模块、音频特征提取与音视频融合模块。实验结果表明,该模型在所提数据集上取得最优性能。

原文摘要 · Abstract (English)

In response to the rising prominence of the Metaverse, omnidirectional videos (ODVs) have garnered notable interest, gradually shifting from professional-generated content (PGC) to user-generated content (UGC). However, the study of audio-visual quality assessment (AVQA) within ODVs remains limited. To address this, we construct a dataset of UGC omnidirectional audio and video (A/V) content. The videos are captured by five individuals using two different types of omnidirectional cameras, shooting 300 videos covering 10 different scene types. A subjective AVQA experiment is conducted on the dataset to obtain the Mean Opinion Scores (MOSs) of the A/V sequences. After that, to facilitate the development of UGC-ODV AVQA fields, we construct an effective AVQA baseline model on the proposed dataset, of which the baseline model consists of video feature extraction module, audio feature extraction and audio-visual fusion module. The experimental results demonstrate that our model achieves optimal performance on the proposed dataset.

音画质量全景视频UGC数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。