arXiv:2607.10162eess.AScs.CL2026-07

研究语音模型能否像人一样感知声音与形状的关联,发现其表现不佳。

Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models

论文配图:Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
图 1 · 摘自论文原文
  • 用真实语音数据对比人类与模型对声音感知的判断差异。
  • 模型在听觉判断上与人类偏差大,忽略关键声学特征如频谱倾斜。
  • 问题出在语音表征,而非视觉能力不足,适合关注语音理解的学者。

声音象征性是人类将语音音素与感知特性(如圆润或尖锐)关联的倾向,主要源于语音的声学特性而非拼写。现有对语音语言模型(SLMs)的研究多依赖文本或图像,而未使用真实语音。本文采用真实人类语音录音,从听觉、跨模态和视觉三个层面比较模型与人类判断的差异。结果表明,SLMs的听觉判断与人类感知对齐较差,且未能捕捉驱动人类直觉的关键声学线索(如频谱倾斜),即使开放权重模型也难以将听到的声音与其对应形状可靠关联。通过仅用视觉的对照实验排除形状感知影响后,问题被定位为语音表征缺陷,说明感知对齐的关键不在于更强的视觉能力,而在于语音表示能否捕捉人类所听的声学线索。

原文摘要 · Abstract (English)

Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling. Whether Speech Language Models (SLMs) share this tendency remains open, as prior evaluations rely on text or images rather than real speech. We study it using genuine human speech recordings, comparing model judgments against human data across the auditory, crossmodal, and visual components of the effect. We find that SLMs' auditory judgments align poorly with human perception and miss the acoustic cues, such as spectral tilt, that drive human intuitions, and open-weight models cannot reliably link a heard sound to its corresponding shape. With a visual-only control ruling out shape perception, the weakness localizes to how speech is represented, suggesting that perceptual alignment depends not on stronger vision but on speech representations that capture the cues humans hear.

语音模型声音象征性感知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。