发现视觉语音模型能识别声音与图像间的非任意关联
Measuring Sound Symbolism in Audio-visual Models
- 构建合成音画数据集,零样本测试模型对声音象征性的感知
- 训练于语音数据的模型表现出显著的声音象征性关联
- 揭示机器学习模型与人类语言认知的潜在共性
近期,视听预训练模型受到广泛关注,并在多种视听任务中表现出色。本研究探究这些模型是否具备声音与视觉表征之间的非任意关联——即声音象征性,这种现象在人类中已被观察到。我们构建了一个包含合成图像与音频样本的专用数据集,采用非参数方法在零样本设置下评估模型表现。结果表明,模型输出与已知的声音象征性模式存在显著相关性,尤其在以语音数据训练的模型中更为明显。这说明此类模型能够捕捉类似人类语言处理中的声音-意义联系,为认知架构与机器学习策略提供了新见解。
原文摘要 · Abstract (English)
Audio-visual pre-trained models have gained substantial attention recently and demonstrated superior performance on various audio-visual tasks. This study investigates whether pre-trained audio-visual models demonstrate non-arbitrary associations between sounds and visual representations$\unicode{x2013}$known as sound symbolism$\unicode{x2013}$which is also observed in humans. We developed a specialized dataset with synthesized images and audio samples and assessed these models using a non-parametric approach in a zero-shot setting. Our findings reveal a significant correlation between the models' outputs and established patterns of sound symbolism, particularly in models trained on speech data. These results suggest that such models can capture sound-meaning connections akin to human language processing, providing insights into both cognitive architectures and machine learning strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。