首个面向语音角色扮演模型的综合性评估基准,填补了语音角色一致性评测空白。
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
- 构建自动化的两阶段流水线,对电影音频与剧本对齐并生成多维角色档案
- 包含1228个角色、65.6小时语音、13335轮对话,覆盖261部电影
- 首次系统评估语音角色扮演模型在长期角色一致性上的表现,适合研究对话系统与角色建模者
近年来大语言模型(LLMs)的进展极大推动了角色扮演对话代理(RPCAs)的发展,这类系统旨在通过一致的角色设定创造沉浸式体验。然而,现有研究存在双重局限:一是主要关注文本模态,忽视了语调、韵律、节奏等关键语音副语言特征,这些特征对表达角色情感和塑造生动形象至关重要;二是语音角色扮演领域长期缺乏标准化评估基准,现有对话数据集仅用于基础能力测试,角色设定模糊不清,难以有效衡量模型在长期角色一致性等核心能力上的表现。为解决这一关键缺口,我们提出VoxRole,首个专为语音角色扮演代理设计的综合性评估基准。该基准包含13335轮多轮对话,总计65.6小时语音,涵盖来自261部电影的1228个独特角色。为构建该资源,我们提出一种新型两阶段自动化流水线:首先将电影音频与剧本对齐,随后利用大语言模型系统性地为每个角色构建多维度档案。基于VoxRole,我们对当前主流语音对话模型进行了多维度评估,揭示了其在保持角色一致性方面的关键优劣。
原文摘要 · Abstract (English)
Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create immersive user experiences through consistent persona adoption. However, current RPCA research faces dual limitations. First, existing work predominantly focuses on the textual modality, entirely overlooking critical paralinguistic features including intonation, prosody, and rhythm in speech, which are essential for conveying character emotions and shaping vivid identities. Second, the speech-based role-playing domain suffers from a long-standing lack of standardized evaluation benchmarks. Most current spoken dialogue datasets target only fundamental capability assessments, featuring thinly sketched or ill-defined character profiles. Consequently, they fail to effectively quantify model performance on core competencies like long-term persona consistency. To address this critical gap, we introduce VoxRole, the first comprehensive benchmark specifically designed for the evaluation of speech-based RPCAs. The benchmark comprises 13335 multi-turn dialogues, totaling 65.6 hours of speech from 1228 unique characters across 261 movies. To construct this resource, we propose a novel two-stage automated pipeline that first aligns movie audio with scripts and subsequently employs an LLM to systematically build multi-dimensional profiles for each character. Leveraging VoxRole, we conduct a multi-dimensional evaluation of contemporary spoken dialogue models, revealing crucial insights into their respective strengths and limitations in maintaining persona consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。