分析不同音频模型对最小有效信号的识别一致性,发现音乐分类转移率仅26%。
If It's Good Enough for You, It's Good Enough for Me: Transferability of Audio Sufficiencies across Models
- 用最小有效信号测试模型间识别一致性,评估信息处理相似性。
- 音乐类型分类的信号转移率约26%,深度伪造检测差异更大。
- 发现模型间信息处理本质差异,适合研究模型可解释性与鲁棒性。
为深入理解不同音频分类模型的信息处理特性,本文提出可迁移性分析:给定一个模型 $f$ 下完成分类的最小充分信号,检验其他模型是否也能对该信号做出相同分类。定义了充分信号可迁移的标准,并在三个任务(音乐流派、情感识别、深度伪造检测)上展开大规模研究。结果显示,音乐流派分类的信号可迁移率约为26%,而其他任务呈现更高变异性;尤其在深度伪造检测中,部分模型表现出显著不同的迁移行为,称之为“扁平地球”模型。进一步分析表明,该方法能揭示传统指标(如准确率、精确率)无法捕捉到的信息论差异。
原文摘要 · Abstract (English)
In order to gain fresh insights about the information processing characteristics of different audio classification models, we propose transferability analysis. Given a minimal, sufficient signal for a classification on a model $f$, transferability analysis asks whether other models accept this minimal signal as having the same classification as it did on $f$. We define what it means for a sufficient signal to be transferable and perform a large study over $3$ different classification tasks: music genre, emotion recognition and deepfake detection. We find that transferability rates vary depending on the task, with sufficient signals for music genre being transferable $\approx26\%$ of the time. The other tasks reveal much higher variance in transferability and reveal that some models, in particular on deepfake detection, have different transferability behavior. We call these models `flat-earther' models. We investigate deepfake audio in more depth, and show that transferability analysis also allows to us to discover information theoretic differences between the models which are not captured by the more familiar metrics of accuracy and precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。