arXiv:2603.12565cs.SDcs.CL2026-03

让日语语音大模型输出更自然口语化内容

Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization

  • 用偏好优化方法调整日语语音大模型输出风格
  • 在新构建的听觉评测集上性能提升显著
  • 适合研究日语语音对话系统的学者参考

语音大模型通常结合语音识别训练的编码器与文本型大语言模型,导致输出习惯书面语风格,不适用于语音合成。这一问题在日语中尤为突出,因其口语与书面语在敬语、句尾助词和句法复杂度上差异显著。本文提出基于偏好的对齐方法,使日语语音大模型输出更简洁、口语化,适合自然语音合成。为严格评估该任务,我们构建了SpokenElyza基准,基于ELYZA-tasks-100并由母语专家进行听觉验证。实验表明,该方法在SpokenElyza上表现显著提升,同时基本保持原书面风格任务性能。我们将公开SpokenElyza以支持未来日语口语对话系统研究。

原文摘要 · Abstract (English)

SpeechLLMs typically combine ASR-trained encoders with text-based LLM backbones, leading them to inherit written-style output patterns unsuitable for text-to-speech synthesis. This mismatch is particularly pronounced in Japanese, where spoken and written registers differ substantially in politeness markers, sentence-final particles, and syntactic complexity. We propose a preference-based alignment approach to adapt Japanese SpeechLLMs for speech-worthy outputs: text that is concise, conversational, and readily synthesized as natural speech. To rigorously evaluate this task, we introduce SpokenElyza, a benchmark for Japanese speech-worthiness derived from ELYZA-tasks-100 with auditory verification by native experts. Experiments show that our approach achieves substantial improvement on SpokenElyza while largely preserving performance on the original written-style evaluation. We will release SpokenElyza to support future research on Japanese spoken dialog systems.

语音生成日语大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。