arXiv:2608.28970cs.SDcs.AI2026-08中稿 · EMNLP被引 1

用AudioLLM诊断语音缺陷,再精准修正,提升语音合成质量。

Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction

论文配图:Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction
图 1 · 摘自论文原文
  • 通过AudioLLM识别语音中的语调问题并生成修正指令
  • 42K条带标注数据训练出可精准执行指令的重合成模型
  • 适合需要精细控制语音韵律的语音合成应用

现有语音合成系统多采用开环单次生成,常出现局部语调缺陷,如重音错位、不自然停顿或语调平坦等问题,而整体指标难以发现。本文提出LoopTTS,一种由判别器引导的滤-判-修正框架,利用AudioLLM诊断基础语音合成模型输出的缺陷,并生成结构化修正指令;随后,一个细粒度指令跟随式语音合成模型(Refiner)基于初始语音、目标文本和指令进行有指导的表达性重合成。为训练该模型,我们构建了包含42,000个样本的Refiner-DB数据集,提供词级语调弱监督标注。人工评估显示,LoopTTS能有效检测显著感知错误并由Refiner修正,其修复质量优于原始输出及实用的开环重生成基线。该模型在重音与停顿控制等定向语调调整任务中也展现出更强的指令遵循能力。

原文摘要 · Abstract (English)

Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened intonation, that utterance-level metrics often fail to expose. We present LoopTTS, a judge-guided Filter-Judge-Refiner framework for recovering low-quality TTS outputs diagnosed by an AudioLLM. Given an initial utterance from a base TTS model, an AudioLLM Judge identifies salient prosodic issues and generates structured refine instructions; a Refiner, our fine-grained instruction-following TTS model, then performs guided expressive re-synthesis conditioned on the initial utterance, target text, and instruction. To train the Refiner, we construct Refiner-DB, a 42K-example AudioLLM-annotated dataset with word-level prosodic weak supervision. Human evaluation on diagnosed low-quality utterances shows that LoopTTS can detect perceptually salient errors and correct them with the Refiner, outperforming raw generated audio and practical open-loop re-generation baselines in recovery quality. The Refiner also demonstrates stronger instruction-following ability for stress and pause control in targeted prosody modification.

语音合成AudioLLM闭环系统韵律控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。