利用视听自监督学习提升声事件定位检测性能,仅需少量标注数据
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
- 通过视听对比学习,让音频与视觉特征在相同方向上对齐
- 100小时无标注视听数据使错误率从36.4降至34.9
- 适合缺乏大量标注数据的声学场景建模任务
本文研究使用一阶全向体麦克风(FOA)采集的空间音频中声事件定位与检测(SELD)问题。传统方法依赖大量带类别和到达方向(DOA)标注的FOA数据训练深度神经网络,但标注成本高,性能受限。为此,提出一种新型自监督预训练方法:利用大量虚拟现实内容中的空间视听数据,假设声源同时被FOA麦克风与全向摄像头捕获,采用对比学习联合训练音频与视觉编码器,使同一声源在相同方向上的音视频嵌入相互靠近。关键在于,音频嵌入直接从原始信号中按方向提取,而视觉嵌入则从对应方向的局部图像块中分别提取,从而促使音频编码器隐含地学习声音类别与方向信息。在DCASE2022 Task 3数据集(20小时标注数据)上,使用100小时无标注视听数据,可将SELD错误率从36.4降至34.9。
原文摘要 · Abstract (English)
This paper describes sound event localization and detection (SELD) for spatial audio recordings captured by firstorder ambisonics (FOA) microphones. In this task, one may train a deep neural network (DNN) using FOA data annotated with the classes and directions of arrival (DOAs) of sound events. However, the performance of this approach is severely bounded by the amount of annotated data. To overcome this limitation, we propose a novel method of pretraining the feature extraction part of the DNN in a self-supervised manner. We use spatial audio-visual recordings abundantly available as virtual reality contents. Assuming that sound objects are concurrently observed by the FOA microphones and the omni-directional camera, we jointly train audio and visual encoders with contrastive learning such that the audio and visual embeddings of the same recording and DOA are made close. A key feature of our method is that the DOA-wise audio embeddings are jointly extracted from the raw audio data, while the DOA-wise visual embeddings are separately extracted from the local visual crops centered on the corresponding DOA. This encourages the latent features of the audio encoder to represent both the classes and DOAs of sound events. The experiment using the DCASE2022 Task 3 dataset of 20 hours shows non-annotated audio-visual recordings of 100 hours reduced the error score of SELD from 36.4 pts to 34.9 pts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。