arXiv:2511.01261cs.SDcs.AI2025-11被引 4

构建首个可复现的语音角色扮演评估框架,提升真人评分一致性。

Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play

  • 设计双策略评估体系:原型匹配与真实感评估
  • 新模型DRAME-Eval相比零样本模型相关性提升至0.629
  • 提供中英双语标注数据集,支持语音模型评测

角色扮演已成为生成模型的重要测试场景,从文本扩展到多模态交互。将角色扮演引入语音能捕捉语调、情感和表达方式,但也带来新的评估挑战。现有方法多依赖音频大语言模型(ALLMs)作为零样本裁判,但会忽略副语言线索,将多个维度压缩为粗略评分,且使用合成语音参考,无法反映真实角色表现。本文提出Speech-DRAME框架,包含三个部分:(i) Speech-DRAME-EvalBench,一个包含中英双语人工标注数据和训练测试协议的评估基准;(ii) DRAME-Eval,经微调的评估模型,显著优于零样本和少样本的ALLMs;(iii) Speech-DRAME-RoleBench,利用DRAME-Eval作为自动裁判,对比语音基础模型(SFMs)。该框架区分两种互补评估策略:原型评估(自上而下,衡量对角色原型的符合度),真实感评估(自下而上,基于真实人类语音,强调细微角色表现)。相较于零样本ALLM裁判,DRAME-Eval在原型评估中与人类评分的相关性从0.480提升至0.629,在真实感评估中从0.390提升至0.625。通过整合透明的基准资源、建模方法与系统级评估,Speech-DRAME提供了首个全面、可复现的语音角色扮演评估基础。

原文摘要 · Abstract (English)

Role-play has become a key testbed for generative models, expanding from text-only dialogue to multimodal interaction. Extending role-play to speech captures prosody, emotion, and delivery, but also poses new evaluation challenges. Current pipelines often use audio large language models (ALLMs) as zero-shot judges, which miss paralinguistic cues, collapse multiple aspects into coarse scores, and rely on synthetic speech references that fail to reflect real-world roles. We present Speech-DRAME, a unified framework that contributes at three levels: (i) Speech-DRAME-EvalBench, an evaluation benchmark with bilingual human-annotated data and protocols for training and testing speech evaluation models (SEMs), (ii) DRAME-Eval, a fine-tuned evaluation model, which substantially outperforms zero-shot and few-shot ALLMs, and (iii) Speech-DRAME-RoleBench, a speech role-play benchmark that leverages DRAME-Eval as an automatic judge to compare speech foundation models (SFMs). Speech-DRAME distinguishes between two complementary evaluation strategies: Archetype Evaluation, a top-down approach measuring adherence to broad role archetypes, and Realism Evaluation, a bottom-up approach grounded in real human speech that emphasizes nuanced role quality. Compared to zero-shot ALLM judges, DRAME-Eval achieves stronger agreement with human ratings (Pearson correlation from 0.480 to 0.629 in archetypes, and 0.390 to 0.625 in realism). By integrating transparent benchmark resources, modeling approaches, and system-level evaluation, Speech-DRAME provides the first comprehensive, reproducible foundation for assessing spoken role-play.

语音评估角色扮演多模态人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。