arXiv:2506.05984eess.AScs.AI2025-06EMNLP被引 28

用听觉大模型自动评估语音风格,效果接近真人。

Audio-Aware Large Language Models as Judges for Speaking Styles

  • 让能理解音频的LLM担任评委,综合判断语调、节奏等说话风格。
  • Gemini-2.5-pro与人类评分一致性达可比肩水平,验证了有效性。
  • 发现当前语音生成模型在自然对话和风格控制上仍有不足。

音频感知的大语言模型(ALLMs)能够理解音频中的文本与非文本信息。本文探索使用ALLMs作为自动评判者,评估语音生成模型(SLMs)在语音风格指令遵循与角色扮演任务中的表现。评估的说话风格包括情感、音量、语速、重音、音高控制及非语言元素。我们采用四种语音语言模型完成任务,并由人类与ALLMs共同评价其输出结果。对比GPT-4o-audio与Gemini-2.5-pro两种ALLM裁判,发现Gemini与人类评判者的一致性达到可比水平。结果表明,ALLMs具备作为语音风格评估工具的潜力;同时,现有SLMs(包括GPT-4o-audio)在风格控制与自然对话生成方面仍存在改进空间。

原文摘要 · Abstract (English)

Audio-aware large language models (ALLMs) can understand the textual and non-textual information in the audio input. In this paper, we explore using ALLMs as an automatic judge to assess the speaking styles of speeches. We use ALLM judges to evaluate the speeches generated by SLMs on two tasks: voice style instruction following and role-playing. The speaking style we consider includes emotion, volume, speaking pace, word emphasis, pitch control, and non-verbal elements. We use four spoken language models (SLMs) to complete the two tasks and use humans and ALLMs to judge the SLMs' responses. We compare two ALLM judges, GPT-4o-audio and Gemini-2.5-pro, with human evaluation results and show that the agreement between Gemini and human judges is comparable to the agreement between human evaluators. These promising results show that ALLMs can be used as a judge to evaluate SLMs. Our results also reveal that current SLMs, even GPT-4o-audio, still have room for improvement in controlling the speaking style and generating natural dialogues.

语音生成大模型评测音频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。