提出新指标MCLP,量化角色扮演语音风格一致性。
Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability
- 用上下文学习机制计算语音片段的连续概率作为风格度量
- MCLP与人工评估高度一致,在多轮对话中提升风格连贯性
- 适用于需要严格角色契合的交互式语音生成任务
大型音频语言模型(LALMs)已将文本到语音(TTS)拓展至交互式角色扮演场景,这类场景要求高表达性和对角色设定的严格遵循。然而现有模型在多轮对话中难以保持与角色档案和场景描述的一致性。核心瓶颈在于缺乏可量化说话风格的客观指标。为此,我们提出均值延续对数似然(MCLP),作为评估指标与强化学习奖励信号,在基于LALM的角色扮演语音合成(RP-TTS)任务中验证有效。MCLP利用预训练LALMs的上下文学习能力,以转录文本、生成语音及重复转录组成的上下文历史为条件,衡量真实语音标记的似然,作为风格连续性的代理指标。此外,我们将MCLP用于强化学习,增强生成语音与角色指令的风格对齐。为支持该任务,我们构建了一个大规模带有丰富场景与角色标注的RP-TTS数据集。实验表明,MCLP与人工判断的风格一致性高度一致,且作为奖励能有效提升RP-TTS性能,在客观指标与主观评价上均取得稳定提升。代码已公开于https://github.com/y-ren16/MCLP。
原文摘要 · Abstract (English)
Recent advances in Large Audio Language Models (LALMs) have extended Text-to-Speech (TTS) to interactive role-play scenarios, which demand high expressiveness and strict adherence to role-play instructions. However, existing models struggle to maintain stylistic consistency with character profiles and scene descriptions across multi-turn dialogues. A critical bottleneck is the lack of objective metrics for quantifying speaking style. To bridge this gap, we propose Mean Continuation Log-Probability (MCLP) as both an evaluation metric and a reward signal, validated on LALM-based Role-Play TTS (RP-TTS) tasks. MCLP leverages the in-context learning capability of pretrained LALMs to measure the likelihood of ground-truth speech tokens conditioned on a contextual history consisting of the transcript, generated speech, and repeated transcript, serving as a proxy for stylistic continuity. Furthermore, we employ MCLP as a reinforcement learning reward to enhance the style alignment between generated speech and role-play instructions. To support this task, we construct a large-scale RP-TTS dataset with rich scene and character annotations. Experiments demonstrate that MCLP is well aligned with human judgments of stylistic consistency and serves as an effective reward for improving RP-TTS, leading to consistent gains in both objective metrics and subjective evaluations. Our code is publicly available at https://github.com/y-ren16/MCLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。