用自监督嵌入和三元组损失提升生成音频的感知质量评估
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
- 结合BEATs与LSTM,通过三元组损失构建感知相似的嵌入空间
- 在无合成数据训练下实现跨域鲁棒的音频质量评估
- 适合关注生成音频主观评价的语音与音乐研究者
我们提出一种自动多维度感知质量预测系统,用于2025年AudioMOS挑战赛第2赛道。任务是针对文本转语音(TTS)、文本转音频(TTA)和文本转音乐(TTM)生成的音频,预测四个感知评分:制作质量、制作复杂度、内容愉悦度和内容实用性。主要挑战在于自然训练数据与合成评估数据之间的领域偏移。为此,我们结合BEATs——一个预训练的基于Transformer的音频表示模型——与多分支长短期记忆(LSTM)预测器,并采用带缓冲区采样的三元组损失,以感知相似性结构嵌入空间。结果表明,该方法提升了嵌入的区分度和泛化能力,使系统在无需合成训练数据的情况下实现跨域鲁棒的音频质量评估。
原文摘要 · Abstract (English)
We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness--for audio generated by text-to-speech (TTS), text-to-audio (TTA), and text-to-music (TTM) systems. A main challenge is the domain shift between natural training data and synthetic evaluation data. To address this, we combine BEATs, a pretrained transformer-based audio representation model, with a multi-branch long short-term memory (LSTM) predictor and use a triplet loss with buffer-based sampling to structure the embedding space by perceptual similarity. Our results show that this improves embedding discriminability and generalization, enabling domain-robust audio quality assessment without synthetic training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。