评估语音表达是否符合语境,提升有声书与对话系统的真实感。
Evaluating the Expressive Appropriateness of Speech in Rich Contexts

- 基于语篇上下文判断语音表达是否恰当,而非仅测情绪强度。
- 构建首个中文口语情境标注数据集,含15维人类评价维度。
- 融合多模型协作与强化学习,显著超越现有语音评估系统。
语音表达评估仍具挑战性,现有方法主要关注情绪强度,忽视语音在特定语境下的表达适宜性。这一局限阻碍了叙事类和交互式应用(如有声书、对话代理)中语音系统的可靠评估。本文提出CEAEval框架,用于评估语音样本是否在语篇层面的叙述语境下,与其隐含的交际意图相契合。为此,我们构建了首个包含真实中文口语表演的情境丰富语音数据集CEAEval-D,提供叙事描述及十五维人类标注,涵盖表达属性与表达适宜性。进一步开发了CEAEval-M模型,整合知识蒸馏、基于规划的多模态协作、自适应音频注意力偏置与强化学习,实现情境丰富的表达适宜性评估。在人工标注测试集上的实验表明,CEAEval-M显著优于现有语音评估与分析系统。
原文摘要 · Abstract (English)
Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its contextual setting. This limitation hinders reliable evaluation of speech systems used in narrative-driven and interactive applications, such as audiobooks and conversational agents. We introduce CEAEval, a Context-rich framework for Evaluating Expressive Appropriateness in speech, which assesses whether a speech sample expressively aligns with the underlying communicative intent implied by its discourse-level narrative context. To support this task, we construct CEAEval-D, the first context-rich speech dataset with real human performances in Mandarin conversational speech, providing narrative descriptions together with fifteen dimensions of human annotations covering expressive attributes and expressive appropriateness. We further develop CEAEval-M, a model that integrates knowledge distillation, planner-based multi-model collaboration, adaptive audio attention bias, and reinforcement learning to perform context-rich expressive appropriateness evaluation. Experiments on a human-annotated test set demonstrate that CEAEval-M substantially outperforms existing speech evaluation and analysis systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。