新基准测试发现主流音频模型听不准音高,可靠性堪忧。
PitchBench: Measuring Pitch Hearing in Audio-Language Models

- 设计28项实验,系统评估模型在不同声学条件下的音高感知能力
- 模型在合成与真实乐器上表现均差,音高识别准确率波动剧烈
- 适合关注音乐理解、多模态感知的AI研究者使用
音频-语言模型(ALMs)在音乐辅导、转录、推荐和制作等现实应用中日益重要,其需具备从声音中可靠推理的能力。然而,现有评测很少直接衡量最基础的音高感知能力。当前方法多通过高阶任务间接评估,且常采用选择题形式,难以判断模型对不同乐器、声学条件和响应格式下细微音高的识别稳定性。本文提出PitchBench,一个涵盖绝对与相对音高感知的系统性评估套件,包含28项实验,覆盖音符持续时间、响度、音源、时间拉伸、背景噪声等变量。任务从孤立音高识别到四声部旋律追踪。评估前沿ALMs发现,其音高感知极不可靠:准确率随音源、持续时间和记谱格式变化显著,即使在受控的合成与乐器信号上也未见稳定表现。研究同时发布PitchBench Python工具包,含数据与生成工具,以支持后续音高感知建模研究。
原文摘要 · Abstract (English)
Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription to captioning, recommendation systems, and music production. More broadly, they are becoming an important component of multimodal AI systems that must reason from sensory input rather than text alone. This makes reliable musical perception a critical prerequisite: if a model cannot accurately hear the structure of sound, it cannot be trusted to reason about, teach, transcribe, or act on audio in the real world. Yet existing benchmarks rarely assess one of the most fundamental musical abilities underlying such perception: pitch hearing. Current evaluations tend to probe pitch hearing only indirectly, through higher-level tasks and often in multiple-choice formats, leaving open how reliably ALMs identify fine-grained pitch across instruments, acoustic conditions, and response formats. We introduce PitchBench, an evaluation suite that systematically measures pitch hearing in ALMs. PitchBench comprises 28 experiments spanning absolute and relative pitch perception within sequences and chords, while varying loudness, note duration, sound source, time stretching, background noise, and other acoustic conditions. Tasks range from identifying individual pitches in isolation to tracking a melodic line within a four-part musical texture. Evaluating frontier ALMs, we find that pitch hearing remains highly unreliable: models perform consistently poorly across settings, with accuracy varying sharply by sound source, note duration, and notation format. Current ALMs do not yet possess stable pitch perception, even for controlled synthetic and instrumental stimuli. Alongside the benchmark, we release PitchBench as a Python package containing the evaluation data and data generation tools to support future work on pitch-aware audio-language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。