arXiv:2511.20972cs.SD2025-11

让对话系统唱歌回应,提升角色扮演互动体验

SingingSDS: A Singing-Capable Spoken Dialogue System for Conversational Roleplay Applications

  • 采用ASR-LLM-SVS模块化流程,实现歌声响应
  • 支持多种角色、音色与旋律配置,适配不同场景
  • 开源可定制,适合互动娱乐与角色扮演应用

随着自动语音识别(ASR)、大语言模型(LLM)和文本到语音(TTS)技术的发展,语音对话系统(SDS)已广泛可用。然而,现有系统大多仅支持常规说话回复。本文提出SingingSDS,一种级联式语音对话系统,通过歌唱而非说话进行回应,在基于角色的扮演和互动娱乐场景中带来更富情感、更易记忆、更愉悦的交互体验。SingingSDS采用模块化的ASR-LLM-SVS流水线,支持多种配置组合:角色人设、ASR与LLM后端、合成语音声学模型(SVS)、旋律来源及音色档案,可根据延迟、质量与音乐风格需求灵活调整。系统提供即插即用的网页演示,代码开源,支持自定义与扩展。演示地址:https://huggingface.co/spaces/espnet/SingingSDS;代码仓库:https://github.com/SingingSDS/SingingSDS。

原文摘要 · Abstract (English)

With recent advances in automatic speech recognition (ASR), large language models (LLMs), and text-to-speech (TTS) technologies, spoken dialogue systems (SDS) have become widely accessible. However, most existing SDS are limited to conventional spoken responses. We present SingingSDS, a cascaded SDS that responds through singing rather than speaking, fostering more affective, memorable, and pleasurable interactions in character-based roleplay and interactive entertainment scenarios. SingingSDS employs a modular ASR-LLM-SVS pipeline and supports a wide range of configurations across character personas, ASR and LLM backends, SVS models, melody sources, and voice profiles, tailored to different needs in terms of latency, quality, and musical style. SingingSDS is available as a plug-and-play web demo, featuring modular, open-source code that supports customization and extension. Demo: https://huggingface.co/spaces/espnet/SingingSDS. Code: https://github.com/SingingSDS/SingingSDS.

语音对话歌声生成角色扮演开源系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。