arXiv:2604.19055cs.SD2026-04中稿 · ACM ICMR 2026被引 1

让虚拟角色说话时既保持人设又自然表达情绪

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis

论文配图:ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis
图 1 · 摘自论文原文
  • 用双轨架构分离音色与语调,分别控制
  • 零样本语音验证错误率仅0.04,情感表达更丰富
  • 适合动漫角色、数字人等需要稳定人设的场景

高保真角色语音合成是沉浸式多媒体应用的核心,尤其在与动漫形象和数字人互动时。然而现有系统难以在不同情绪下保持一致的角色特征。为此,我们提出ATRIE,采用人物-语调双轨(P2-DT)架构,将生成过程解耦为静态音色轨(通过标量量化)和动态语调轨(通过分层流匹配),后者由140亿参数大语言模型教师蒸馏而来。该设计实现了强身份保持(零样本语音验证错误率EER: 0.04)与丰富情感表达。在扩展的AnimeTTS-Bench(50个角色)上评估,ATRIE在生成质量和跨模态检索上均达领先水平(mAP: 0.75),确立了人物驱动多媒体内容创作的新范式。

原文摘要 · Abstract (English)

High-fidelity character voice synthesis is a cornerstone of immersive multimedia applications, particularly for interacting with anime avatars and digital humans. However, existing systems struggle to maintain consistent persona traits across diverse emotional contexts. To bridge this gap, we present ATRIE, a unified framework utilizing a Persona-Prosody Dual-Track (P2-DT) architecture. Our system disentangles generation into a static Timbre Track (via Scalar Quantization) and a dynamic Prosody Track (via Hierarchical Flow-Matching), distilled from a 14B LLM teacher. This design enables robust identity preservation (Zero-Shot Speaker Verification EER: 0.04) and rich emotional expression. Evaluated on our extended AnimeTTS-Bench (50 characters), ATRIE achieves state-of-the-art performance in both generation and cross-modal retrieval (mAP: 0.75), establishing a new paradigm for persona-driven multimedia content creation.

语音合成角色声音情感表达数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。