提出可自动评估音频分离效果的多模态工具,更贴近人耳感知。
SAM Audio Judge: A Unified Multimodal Framework for Perceptual Evaluation of Audio Separation
- 基于文本、视觉、片段三种提示,实现无参考的细粒度评估
- 在语音、音乐、通用声学场景中均与人类判断高度一致
- 适合用于数据清洗、伪标签生成和模型排序
音频分离的性能评估仍是复杂挑战,现有指标常与人类感知不一致,且依赖真实参考信号。主观听觉测试虽为金标准,但成本高、难扩展。本文提出SAM Audio Judge(SAJ),一种多模态、细粒度、无参考的客观评价指标,与人类感知高度一致。SAJ覆盖语音、音乐、通用声音事件三大音频域,支持文本、视觉、片段三种提示输入,涵盖召回率、精确率、忠实度和整体评价四个维度。该方法还可应用于数据过滤、大规模数据集伪标签生成及音频分离模型的重排序。代码与预训练模型已开源:https://github.com/facebookresearch/sam-audio。
原文摘要 · Abstract (English)
The performance evaluation remains a complex challenge in audio separation, and existing evaluation metrics are often misaligned with human perception, course-grained, relying on ground truth signals. On the other hand, subjective listening tests remain the gold standard for real-world evaluation, but they are expensive, time-consuming, and difficult to scale. This paper addresses the growing need for automated systems capable of evaluating audio separation without human intervention. The proposed evaluation metric, SAM Audio Judge (SAJ), is a multimodal fine-grained reference-free objective metric, which shows highly alignment with human perceptions. SAJ supports three audio domains (speech, music and general sound events) and three prompt inputs (text, visual and span), covering four different dimensions of evaluation (recall, percision, faithfulness, and overall). SAM Audio Judge also shows potential applications in data filtering, pseudo-labeling large datasets and reranking in audio separation models. We release our code and pre-trained models at: https://github.com/facebookresearch/sam-audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。