arXiv:2412.17667cs.SDcs.MM2024-12NAACL被引 67

一个统一的语音音频音乐评估工具,支持65种指标

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

  • 提供灵活配置的Python接口,支持多种评估场景
  • 包含65个指标,共729种变体,覆盖多类参考数据
  • 适合语音、音频、音乐生成与处理任务的评估研究

本文提出VERSA,一个统一且标准化的语音、音频和音乐信号评估工具包。该工具包采用面向Python的接口设计,具备灵活配置与依赖管理能力,使用便捷高效。完整安装后,VERSA提供65个评估指标,涵盖729种不同配置下的指标变体,可基于匹配或非匹配参考音频、文本转录和文本描述等外部资源进行评估。作为轻量但全面的工具包,VERSA适用于多种下游应用场景的评估。文中展示了其在音频编码、语音合成、语音增强、歌唱合成及音乐生成等任务中的实际应用案例。工具开源地址:https://github.com/wavlab-speech/versa。

原文摘要 · Abstract (English)

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA offers 65 metrics with 729 metric variations based on different configurations. These metrics encompass evaluations utilizing diverse external resources, including matching and non-matching reference audio, text transcriptions, and text captions. As a lightweight yet comprehensive toolkit, VERSA is versatile to support the evaluation of a wide range of downstream scenarios. To demonstrate its capabilities, this work highlights example use cases for VERSA, including audio coding, speech synthesis, speech enhancement, singing synthesis, and music generation. The toolkit is available at https://github.com/wavlab-speech/versa.

音频评估语音合成音乐生成工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。