用自然语言描述构建语音质量评估新数据集,提升评测的可解释性。
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
- 基于11个关键维度收集语音质量的自然语言描述与推理
- 微调后的听觉大模型能准确识别噪声类型和时间特征
- 适合研究语音理解、人机交互与可解释评估的学者
本文通过引入自然语言描述,探索语音质量评估的新视角,相较于传统数值评分提供更丰富、更细致的洞察。自然语言反馈包含指导性建议和详细评价,但现有数据集缺乏支持该方法的全面标注。为此,我们提出QualiSpeech,一个涵盖11个关键维度的低层级语音质量评估数据集,包含详细的自然语言评论,涵盖推理与上下文信息。同时,我们构建了QualiSpeech基准,用于评估听觉大语言模型(auditory LLMs)在低层级语音理解方面的能力。实验表明,微调后的听觉大模型能可靠生成关于噪声和失真的详细描述,有效识别其类型与时间特性。结果进一步凸显了引入推理机制对提升评估准确性与可靠性的重要潜力。数据集将发布于https://huggingface.co/datasets/tsinghua-ee/QualiSpeech。
原文摘要 · Abstract (English)
This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset will be released at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。