FAD评估结果受编码器任务影响,该研究揭示其内在偏差并提出更公平的评估方法。
An Empirical Analysis of Task-Induced Encoder Bias in Fréchet Audio Distance
- 将FAD拆解为召回、精度与对齐(语义/结构),用对数归一化实现跨编码器公平比较
- 六种编码器在两个数据集上测试,发现重建、语音识别、分类任务各有侧重且存在四维权衡
- 指出无通用评估编码器,未来应构建与人类感知一致的原生评估模型
Fréchet Audio Distance (FAD) 是文本到音频生成的主流评估指标,但其得分受底层编码器嵌入空间的影响。编码器的训练任务决定了哪些声学特征被保留或丢弃,导致 FAD 继承系统性任务偏差。本文将评估分解为召回、精度与对齐(细分为语义和结构维度),采用对数尺度归一化以实现跨编码器的公平比较。在两个数据集上对六种编码器进行受控实验,揭示出一种四轴权衡:基于重建的 AudioMAE 在精度敏感性上表现突出;ASR 训练的 Whisper 在结构检测上占优,但对信号退化不敏感;分类训练的 VGGish 在语义检测上最强,却惩罚了类内合法变化。由于不存在单一通用评估编码器,未来的度量需转向与人类感知内在对齐的原生评估编码器。
原文摘要 · Abstract (English)
Fréchet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or discarded, causing FAD to inherit systematic task-induced biases. We decompose evaluation into Recall, Precision, and Alignment (split into semantic and structural dimensions), using log-scale normalization for fair cross-encoder comparison. Controlled experiments on six encoders across two datasets reveal a four-axis trade-off: reconstruction-based AudioMAE leads precision sensitivity; ASR-trained Whisper dominates structural detection but is blind to signal degradation; classification-trained VGGish maximizes semantic detection but penalizes legitimate intra-class variation. Since no single encoder is a universal evaluator, future metrics must shift toward evaluation-native encoders intrinsically aligned with human perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。