构建首个大规模音频角色扮演数据集,助力大模型精准模仿人物音色与性格。
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
- 从13部剧集收集超1000小时音频,标注100万+对话的声线与语境信息
- 新模型在语音个性化上达0.31分,内容个性化提升38%,超越多个基线模型
- 提供评估框架和6个角色模型,适合研究声音驱动的角色扮演任务
高质量多模态数据集对提升大语言模型的角色扮演能力至关重要。现有工作多聚焦文本角色模拟,而音频角色扮演(ARP)因需语义内容与声线特征同步对齐,面临独特挑战。为此,我们提出AudioRole,一个源自13部电视剧、覆盖1000小时以上、包含100万+角色相关对话的精心构建数据集,提供带说话人身份与上下文元数据标注的音视频同步对齐样本。为验证数据集有效性,我们引入ARP-Eval,一种双维度评估框架,衡量回复质量与角色忠实度。实证显示,在AudioRole上训练的GLM-4-Voice(称作ARP-Model)平均声线个性化得分达0.31,显著优于原始GLM-4-Voice及更强的MiniCPM-O-2.6(支持单次角色扮演),内容个性化得分为0.36,较未训练原模型提升约38%,与MiniCPM-O-2.6持平。AudioRole涵盖115位以上主要角色,支持6个已训练的ARP-Model进行不同角色扮演,并提供完整评估协议。该资源为推进基于音频的角色扮演研究提供了关键支持。
原文摘要 · Abstract (English)
The creation of high-quality multimodal datasets remains fundamental for advancing role-playing capabilities in large language models (LLMs). While existing works predominantly focus on text-based persona simulation, Audio Role-Playing (ARP) presents unique challenges due to the need for synchronized alignment of semantic content and vocal characteristics. To address this gap, we propose AudioRole, a meticulously curated dataset from 13 TV series spanning 1K+ hours with 1M+ character-grounded dialogues, providing synchronized audio-text pairs annotated with speaker identities and contextual metadata. In addition, to demonstrate the effectiveness of the dataset, we introduced ARP-Eval, a dual-aspect evaluation framework that assesses both response quality and role fidelity. Empirical validation showing GLM-4-Voice trained on AudioRole (which we called ARP-Model) achieve an average Acoustic Personalization score of 0.31, significantly outperforming the original GLM-4-voice and the more powerful model MiniCPM-O-2.6, which specifically supports role-playing in one-shot scenarios. The ARP-Model also achieves a Content Personalization score of 0.36, surpassing the untrained original model by about 38% and maintaining the same level as MiniCPM-O-2.6. AudioRole features dialogues from over 115 main characters, 6 trained ARP-Models that role-play different characters, and evaluation protocols. Together, they provide an essential resource for advancing audio-grounded role-playing research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。