重构VGGSound数据集,精准评估视听模型能力
VGGSounder: Audio-Visual Evaluations for Foundation Models
- 重新标注多标签数据,解决原数据错标与模态错位问题
- 新指标发现加模态后模型性能下降,暴露融合缺陷
- 适合研究多模态模型评估与鲁棒性提升的学者
视听基础模型的兴起凸显了可靠评估其多模态理解能力的重要性。当前广泛使用的VGGSound数据集存在标注不全、类别部分重叠及模态对齐错误等问题,导致对听觉与视觉能力的评估失真。为此,我们提出VGGSounder,一个全面重新标注的多标签测试集,扩展了VGGSound,专为评估视听基础模型而设计。VGGSounder包含详细的模态标注,支持对各模态性能的精确分析。此外,通过新提出的模态混淆度量,我们揭示了在增加另一输入模态时模型性能下降的现象,暴露出现有模型在跨模态融合中的局限性。
原文摘要 · Abstract (English)
The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。