评测语音模型能否听懂口头指令调整说话风格,推动更自然的人机对话。
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
- 提出语音风格适应新任务,通过口语指令改变音色、语调和角色。
- 构建中英双语基准VStyle,覆盖四类语音生成场景,含真实对话数据。
- 设计自动评估框架,客观衡量语音忠实度、风格匹配与自然度。
语音语言模型(SLMs)已成为语音理解与生成的统一范式,支持自然的人机交互。然而,当前研究多关注语义准确性和指令遵循,对模型根据口语指令调整说话风格的能力关注不足。本文提出语音风格适应(VSA)任务,考察SLMs是否能根据自然语言口令改变音色、语调或角色。为此,我们构建了VStyle——一个涵盖声学属性、自然语言指令、角色扮演和隐含共情四类场景的中英双语基准。同时提出大型音频语言模型作为评判者(LALM as a Judge)的评估框架,分阶段评估输出在文本忠实度、风格契合度和自然度上的表现,确保评估可复现且客观。在商用系统与开源模型上的实验表明,现有模型在可控风格适配方面仍存在明显局限,凸显该任务的新颖性与挑战性。我们公开发布VStyle及评估工具包,旨在为推进以人为本的语音交互提供基础。数据集与代码已开放:https://junzhan2000.github.io/VStyle.github.io/
原文摘要 · Abstract (English)
Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and instruction following, the ability of SLMs to adapt their speaking style based on spoken instructions has received limited attention. We introduce Voice Style Adaptation (VSA), a new task that examines whether SLMs can modify their speaking style, such as timbre, prosody, or persona following natural language spoken commands. To study this task, we present VStyle, a bilingual (Chinese & English) benchmark covering four categories of speech generation: acoustic attributes, natural language instruction, role play, and implicit empathy. We also introduce the Large Audio Language Model as a Judge (LALM as a Judge) framework, which progressively evaluates outputs along textual faithfulness, style adherence, and naturalness, ensuring reproducible and objective assessment. Experiments on commercial systems and open source SLMs demonstrate that current models face clear limitations in controllable style adaptation, highlighting both the novelty and challenge of this task. By releasing VStyle and its evaluation toolkit, we aim to provide the community with a foundation for advancing human centered spoken interaction. The dataset and code are publicly available at \href{https://junzhan2000.github.io/VStyle.github.io/}{project's homepage}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。