让日语语音大模型输出更自然口语化内容
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
- 用偏好优化方法调整日语语音大模型输出风格
- 在新构建的听觉评测集上性能提升显著
- 适合研究日语语音对话系统的学者参考
语音大模型通常结合语音识别训练的编码器与文本型大语言模型,导致输出习惯书面语风格,不适用于语音合成。这一问题在日语中尤为突出,因其口语与书面语在敬语、句尾助词和句法复杂度上差异显著。本文提出基于偏好的对齐方法,使日语语音大模型输出更简洁、口语化,适合自然语音合成。为严格评估该任务,我们构建了SpokenElyza基准,基于ELYZA-tasks-100并由母语专家进行听觉验证。实验表明,该方法在SpokenElyza上表现显著提升,同时基本保持原书面风格任务性能。我们将公开SpokenElyza以支持未来日语口语对话系统研究。
原文摘要 · Abstract (English)
SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This mismatch is particularly pronounced in Japanese, where spoken and written registers differ substantially in politeness markers, sentence-final particles, and syntactic complexity. We propose a preference-based alignment approach to adapt Japanese SpeechLLMs for speech-worthy outputs: text that is concise, conversational, and readily synthesized as natural speech. To rigorously evaluate this task, we introduce SpokenElyza, a benchmark for Japanese speech-worthiness derived from ELYZA-tasks-100 with auditory verification by native experts. Experiments show that our approach achieves substantial improvement on SpokenElyza while largely preserving performance on the original written-style evaluation. We will release SpokenElyza to support future research on Japanese spoken dialog systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。