评测语音大模型在对话中控制语调强度的能力。
StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control
- 构建多轮对话基准,评估情绪、语速、音量、音高四维风格控制。
- 发现主流语音模型与通用语言模型在风格控制上存在明显差距。
- 适合语音合成、人机交互研究者参考,推动个性化语音对话发展。
语音语言模型(SLMs)通过引入副语言信息,显著提升了文本大模型的交互能力。为实现更自然的个性化对话体验,当前SLMs已能根据用户提示在对话中解析并控制说话风格强度。然而,缺乏系统性基准来量化评估对话中的风格强度控制能力。本文提出StyleBench,一个面向多轮对话的基准,全面评估情绪、语速、音量和音高四个维度的风格控制能力。实验结果揭示了领先SLMs与通用语言模型(OLMs)之间的性能差距,揭示了潜在原因,并为未来研究指明方向。
原文摘要 · Abstract (English)
Speech language models (SLMs) have significantly extended the interactive capability of text-based Large Language Models (LLMs) by incorporating paralinguistic information. For more realistic interactive experience with customized styles, current SLMs have managed to interpret and control speaking style intensity from user prompts during the dialogue process. However, there remains a lack of systematic benchmarks that quantifies and evaluates the style intensity control ability in conversations. In this paper, we propose StyleBench, a multi-turn dialogue benchmark for comprehensively evaluating the style intensity control ability across four dimensions: emotion, speed, volume, and pitch. Our results reveal the performance gaps between leading SLMs and omni language models (OLMs), suggesting the underlying reasons and promising approaches for future exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。