用声音和非声音数据联合训练,提升鸟类识别准确率
Combining Audio and Non-Audio Inputs in Evolved Neural Networks for Ovenbird
- 将声谱图与栖息地、物候等非音频数据融合输入网络
- 融合多源信息的模型准确率显著高于单一输入模型
- 适合做生物多样性监测或生态数据融合研究者参考
近年来,神经网络在自动化物种分类中的应用日益广泛,尤其图像分类中卷积神经网络(CNN)表现优异。对于音频数据,基于CNN的识别器常通过声谱图进行物种分类。当前方法多仅使用声谱图作为输入,但实际研究中还拥有物种栖息地偏好、物候及分布范围等非音频数据,这些信息可能提升分类性能。本文提出将非音频数据与声谱图共同输入单物种识别神经网络,以提升分类准确率,并验证了性能提升是否源于参数量增加。结果表明,融合双源输入的网络在相似参数规模下,分类准确率显著优于仅使用单一输入的模型。
原文摘要 · Abstract (English)
In the last several years the use of neural networks as tools to automate species classification from digital data has increased. This has been due in part to the high classification accuracy of image classification through Convolutional Neural Networks (CNN). In the case of audio data CNN based recognizers are used to automate the classification of species in audio recordings by using information from sound visualization (i.e., spectrograms). It is common for these recognizers to use the spectrogram as their sole input. However, researchers have other non-audio data, such as habitat preferences of a species, phenology, and range information, available that could improve species classification. In this paper we present how a single-species recognizer neural network's accuracy can be improved by using non-audio data as inputs in addition to spectrogram information. We also analyze if the improvements are merely a result of having a neural network with a higher number of parameters instead of combining the two inputs. We find that networks that use the two different inputs have a higher classification accuracy than networks of similar size that use only one of the inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。