让大模型角色扮演更自然,通过推理时干预实现思维与声音同步。
DeSRPA: Decoupled Speech Role-Playing Agent via Inference-Time Intervention

- 推理时干预冻结的模型,分离思维与表达控制。
- 在两个基准上显著提升角色性格与情绪一致性。
- 无需训练,适合快速部署到新角色且语音自然度高。
尽管大语言模型(LLMs)已革新文本角色扮演,但构建沉浸式语音角色扮演代理(SRPAs)仍需在认知推理与副语言特征间建立无缝连接。现有SRPAs主要依赖端到端(E2E)微调,但该方法因依赖特定角色数据,导致对未见角色泛化能力差,并带来“模态对齐代价”,削弱了原始LLM的推理能力。我们提出DeSRPA,一种基于推理时干预的代理框架,使用冻结骨干网络实现角色扮演。DeSRPA采用双层控制向量机制:内部认知引导与外部表达渲染,实现“心智”与“声线”的同步。在SpeechRole和OmniCharacter基准上的实验表明,DeSRPA在性格与情绪一致性方面显著优于E2E基线。其语音自然度接近专有模型GPT-4o Audio,同时保持可扩展性和零训练特性。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have revolutionized text-based role-playing, creating immersive Speech Role-Playing Agents (SRPAs) requires a seamless bridge between cognitive reasoning and paralinguistic nuances. Current SRPAs primarily rely on end-to-end (E2E) fine-tuning. However, this paradigm suffers from poor generalization to unseen characters due to its reliance on role-specific data, while imposing a "modality alignment tax" that degrades intrinsic LLM reasoning capabilities. We propose DeSRPA, an agentic framework for character role play via inference-time intervention on frozen backbones. DeSRPA employs a dual-level control vector mechanism, Internal Cognitive Steering and External Expressive Rendering, to synchronize "mind" and "voice". Experiments on SpeechRole and OmniCharacter benchmarks demonstrate that DeSRPA significantly outperforms E2E baselines in personality and emotional consistency. It achieves high speech naturalness, narrowing the gap with proprietary models like GPT-4o Audio, while remaining a scalable and training-free paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。