通过众包构建迄今最大最多样化的音视频质量评估数据集
Scaling Audio-Visual Quality Assessment Dataset via Crowdsourcing
- 采用众包实验框架,实现跨环境可靠标注
- 覆盖1620段用户生成音视频,涵盖多种质量与内容场景
- 新增多模态感知标注,支持内容与感知关系研究
音视频质量评估(AVQA)研究受限于现有数据集规模小、内容与质量多样性不足,且仅提供整体评分。为此,我们提出一种实用的AVQA数据集构建方法:首先设计众包主观实验框架,突破实验室限制,实现跨环境可靠标注;其次采用系统化数据准备策略,确保覆盖广泛的质量水平和语义场景;第三,扩展数据集包含额外标注,支持多模态感知机制及其与内容关联的研究。最终通过YT-NTU-AVQ验证该方法,该数据集为迄今最大最多样化的AVQA数据集,包含1,620段用户生成的音视频序列。数据集与平台代码已公开于https://github.com/renyu12/YT-NTU-AVQ。
原文摘要 · Abstract (English)
Audio-visual quality assessment (AVQA) research has been stalled by limitations of existing datasets: they are typically small in scale, with insufficient diversity in content and quality, and annotated only with overall scores. These shortcomings provide limited support for model development and multimodal perception research. We propose a practical approach for AVQA dataset construction. First, we design a crowdsourced subjective experiment framework for AVQA, breaks the constraints of in-lab settings and achieves reliable annotation across varied environments. Second, a systematic data preparation strategy is further employed to ensure broad coverage of both quality levels and semantic scenarios. Third, we extend the dataset with additional annotations, enabling research on multimodal perception mechanisms and their relation to content. Finally, we validate this approach through YT-NTU-AVQ, the largest and most diverse AVQA dataset to date, consisting of 1,620 user-generated audio and video (A/V) sequences. The dataset and platform code are available at https://github.com/renyu12/YT-NTU-AVQ
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。