arXiv:2503.19161eess.AScs.SD2025-03

用视觉模型分析音频音高轮廓,跨领域表现更优。

Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach

  • 用预训练图像模型提取音高轮廓特征,无需传统音高追踪。
  • 在8个跨领域任务中均优于传统方法,尤其在快速变化频段表现突出。
  • 适合研究音乐、语音、生物声学等多领域音高对比的学者。

本研究将音高轮廓视为音乐、语音、生物声学及日常声响等多音频领域的共性语义结构。分析音高轮廓有助于理解音高在听觉感知中的普遍作用,并深化对人类与动物听觉机制的认识。传统音高追踪方法虽针对音乐和语音优化,但在更广频率范围和更快速音高变化的其他音频领域面临挑战。本文提出一种基于视觉的音高轮廓分析方法,无需显式音高追踪。该方法采用预训练于自然图像目标检测的卷积神经网络,通过合成生成的音高轮廓数据集进行微调,从短时音频片段的时间-频率表示中提取关键轮廓参数。选取来自四个音频领域的八项下游任务构成具有挑战性的跨域评估场景。结果表明,所提方法在多项任务中持续超越基于音高追踪的传统技术,说明该视觉方法为跨音频领域音高轮廓特征的比较研究奠定了基础。

原文摘要 · Abstract (English)

This study examines pitch contours as a unifying semantic construct prevalent across various audio domains including music, speech, bioacoustics, and everyday sounds. Analyzing pitch contours offers insights into the universal role of pitch in the perceptual processing of audio signals and contributes to a deeper understanding of auditory mechanisms in both humans and animals. Conventional pitch-tracking methods, while optimized for music and speech, face challenges in handling much broader frequency ranges and more rapid pitch variations found in other audio domains. This study introduces a vision-based approach to pitch contour analysis that eliminates the need for explicit pitch-tracking. The approach uses a convolutional neural network, pre-trained for object detection in natural images and fine-tuned with a dataset of synthetically generated pitch contours, to extract key contour parameters from the time-frequency representation of short audio segments. A diverse set of eight downstream tasks from four audio domains were selected to provide a challenging evaluation scenario for cross-domain pitch contour analysis. The results show that the proposed method consistently surpasses traditional techniques based on pitch-tracking on a wide range of tasks. This suggests that the vision-based approach establishes a foundation for comparative studies of pitch contour characteristics across diverse audio domains.

音高分析跨域学习视觉模型音频表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。