arXiv:2507.08603cs.AIcs.SD2025-07ACL被引 4

用多模型重写提升语音指令数据质量,无需人工标注。

Unlocking Speech Instruction Data Potential with Query Rewriting

  • 用多个大模型融合重写语音指令文本,改善TTS模型适配性。
  • 零样本重写使数据可用率从72%提升至93%。
  • 适合构建大规模语音指令数据集的研究者使用。

端到端大语音语言模型(LSLMs)在响应延迟和语音理解方面展现强大潜力,具备跨任务的通用智能。然而,由于缺乏高质量数据集及训练任务严重偏倚,语音指令执行能力尚未充分实现。以往方法借助大量ASR数据,利用大语言模型(LLMs)续写语音文本以构建语音指令数据集,但因生成结果与真实人类回应存在差距,导致问题被放大。人工收集和标注语音指令数据成本高昂,因此采用语音合成构建大规模数据集成为更优选择。尽管现代文本转语音(TTS)模型已接近人类水平,但对分布外文本指令的语音转换仍受训练数据分布限制。为此,本文提出一种基于多LLM知识融合的查询重写框架,通过多个智能体标注与验证合成语音,实现无需人工标注的高质量语音指令数据集构建。实验表明,该方法可通过零样本重写将文本指令转换为更适配TTS模型的分布,使数据可用率从72%提升至93%,在需复杂知识与上下文理解的任务中更具优势。

原文摘要 · Abstract (English)

End-to-end Large Speech Language Models~(\textbf{LSLMs}) demonstrate strong potential in response latency and speech comprehension capabilities, showcasing general intelligence across speech understanding tasks. However, the ability to follow speech instructions has not been fully realized due to the lack of datasets and heavily biased training tasks. Leveraging the rich ASR datasets, previous approaches have used Large Language Models~(\textbf{LLMs}) to continue the linguistic information of speech to construct speech instruction datasets. Yet, due to the gap between LLM-generated results and real human responses, the continuation methods further amplify these shortcomings. Given the high costs of collecting and annotating speech instruction datasets by humans, using speech synthesis to construct large-scale speech instruction datasets has become a balanced and robust alternative. Although modern Text-To-Speech~(\textbf{TTS}) models have achieved near-human-level synthesis quality, it is challenging to appropriately convert out-of-distribution text instruction to speech due to the limitations of the training data distribution in TTS models. To address this issue, we propose a query rewriting framework with multi-LLM knowledge fusion, employing multiple agents to annotate and validate the synthesized speech, making it possible to construct high-quality speech instruction datasets without relying on human annotation. Experiments show that this method can transform text instructions into distributions more suitable for TTS models for speech synthesis through zero-shot rewriting, increasing data usability from 72\% to 93\%. It also demonstrates unique advantages in rewriting tasks that require complex knowledge and context-related abilities.

语音指令TTS数据构建多模型融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。