让语音对话模型听指令控制说话风格和互动行为。
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models
- 仅微调语言模型,2000小时数据即可训练
- 支持语音、话题、互动行为等多维度指令控制
- 首个开源可复现的全双工可控语音模型
语音对话系统不仅需要准确生成语音,还需动态适应语境以实现自然流畅的交互。现有系统普遍缺乏定制能力,限制了其真实感与可用性。本文提出首个开源、指令驱动的全双工语音对话模型,可在典型学术资源条件下高效训练。通过冻结音频编码器,仅微调语言模型,模型仅需2000小时数据,无需大规模预训练或多阶段优化。模型可遵循显式指令,控制说话人声音、对话主题、互动行为(如回应、打断)及对话发起方式。我们设计了单阶段训练流程,并系统分析了各项设计选择。模型与训练代码均已公开,推动可控全双工语音系统的研究可复现性。
原文摘要 · Abstract (English)
Spoken conversational systems require more than accurate speech generation to have human-like conversations: to feel natural and engaging, they must produce conversational behaviour that adapts dynamically to the context. Current spoken conversational systems, however, rarely allow such customization, limiting their naturalness and usability. In this work, we present the first open, instruction-following full-duplex conversational speech model that can be trained efficiently under typical academic resource constraints. By keeping the audio encoder frozen and finetuning only the language model, our model requires just 2,000 hours of data, without relying on large-scale pretraining or multi-stage optimization. The model can follow explicit instructions to control speaker voice, conversation topic, conversational behaviour (e.g., backchanneling and interruptions), and dialogue initiation. We propose a single-stage training protocol and systematically analyze design choices. Both the model and training code is released to enable reproducible research on controllable full-duplex speech systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。