对比学习的音频嵌入编码什么,取决于语料结构而非数据量。
Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding
- 通过控制语料结构让韵律成为唯一区分信号,提升关键词识别
- 在相同数据量下,带情感标签的挖掘语料无法提升情绪识别性能
- 适合关注对比学习机制、多模态表示设计的研究者
当对比表示缺乏某一属性时,扩充语料是常规做法。本文发现:在特定情况下扩大语料无效。将词汇-语音轮次加入冻结基线多模态嵌入模型,使零样本关键词检测提升76点,但语音情绪识别下降14点。该损失非容量限制:在7,442段受韵律控制的语料上微调,可恢复情绪识别至预语音水平,仅代价5点关键词准确率。同样数据量下,29,428条显式标注情绪的挖掘语料,其情绪性能仅下降0.0007。关键差异在于结构:对比目标仅在批内负样本必须依赖该属性才能分离时才编码它。受控语料固定句子内容,韵律成为唯一区分信号;而挖掘语料虽标注情绪,但仍可由场景内容分离。干预实验证实因果性:提高字幕相似性不能恢复情绪性能,但压缩字幕多样性使情绪成为唯一分离轴后,情绪识别回升8.9点(三组种子),非表演语料也有小幅正向增益,同时关键词准确率回落。语料结构,而非规模或字幕词汇,决定对比音频嵌入所编码的内容。
原文摘要 · Abstract (English)
Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding model raises zero-shot keyword spotting by 76 points while reducing speech-emotion recognition by 14. The loss is not a capacity limit: fine-tuning on 7,442 clips from a prosody-controlled corpus recovers emotion past its pre-speech level at a five-point keyword cost. Nor is it data volume: 29,428 mined clips whose captions explicitly name emotions, at matched exposure, move emotion by -0.0007. The difference is structural: a contrastive objective encodes an attribute only when the in-batch negatives cannot be separated without it; the controlled corpus holds sentence content fixed, so prosody is the only separating signal, whereas mined captions name emotion yet remain separable by scene content. Intervention on the same audio confirms causality: raising caption similarity does not recover emotion, but collapsing caption diversity so that emotion becomes the only separating axis recovers it by 8.9 points across three seeds, with a smaller, same-signed gain on a non-acted corpus, while keyword accuracy trades back. Corpus structure, not size or caption vocabulary, controls what a contrastive audio embedding encodes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。