用语音大模型评估角色扮演的多维匹配度,提升对话真实性。
Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning

- 通过音频大模型构建跨模态角色对齐评估框架
- 在真实与生成语音上实现92.3%的评估准确率
- 适合语音角色扮演系统开发与评测人员
多模态大模型的快速发展推动了语音对话系统中多样角色的模拟,开启了新型交互范式。角色特征不仅体现在文本回复中,也通过语音中的副语言信息体现,而这些信息难以量化,给角色扮演代理的对齐评估带来挑战。为此,我们提出RoleJudge评估框架,利用音频大语言模型系统性地评估语音与角色在多个模态和维度上的对齐程度。同时,我们引入RoleChat,首个包含思维链标注的语音角色扮演评估数据集,包含多样化的真实与大模型生成语音样本。基于该数据集,我们采用多阶段训练范式,并在强化学习中引入标准对齐机制,缓解优化过程中的奖励错位问题。实验结果在准确率与主观评估上均显示RoleJudge优于多种基线模型,验证了多维度评估框架的有效性。
原文摘要 · Abstract (English)
The rapid evolution of multimodal large models has revolutionized the simulation of diverse characters in speech dialogue systems, enabling a novel interactive paradigm. Character attributes are manifested not only in textual responses but also through vocal features, as speech conveys rich paralinguistic information that is challenging to quantify. This poses significant difficulties in evaluating the character alignment of role-playing agents. To address these challenges, we present RoleJudge, an evaluation framework that leverages audio large language models to systematically assess the alignment between speech and character across multiple modalities and dimensions. Furthermore, we introduce RoleChat, the first voice role-playing evaluation dataset enriched with chain-of-thought reasoning annotations, comprising a diverse set of authentic and LLM-generated speech samples. Utilizing this dataset, we implement a multi-stage training paradigm and incorporate Standard Alignment in reinforcement learning to mitigate reward misalignment during optimization. Experimental results in terms of accuracy and subjective assessment demonstrate that RoleJudge outperforms various baseline models, validating the effectiveness of our multidimensional evaluation framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。