arXiv:2609.08204cs.SD2026-09

解决语音合成指令监督的语义漂移问题,提升指令遵循准确率。

Stabilizing Instruction Supervision for Instruct-TTS via Controllable Diversification and Drift Filtering

论文配图:Stabilizing Instruction Supervision for Instruct-TTS via Controllable Diversification and Drift Filtering
图 1 · 摘自论文原文
  • 通过可控指令多样化系统扩展数据,增强覆盖
  • 利用大模型过滤语义漂移,将错误率从40.4%降至15.4%
  • 基于声学扰动对齐属性,提升韵律控制精度

Instruct-TTS 系统通过大语言模型(LLM)重写结构化风格标签生成自然语言指令,但我们发现超过40%的无约束重写存在语义漂移,破坏监督信号并削弱泛化能力。本文将此问题形式化为指令监督不稳定性,并提出一种以数据为中心的稳定化方案,包含三个机制:可控指令多样化用于系统性扩展、基于LLM的漂移过滤用于质量控制、属性对齐的监督机制将韵律控制与声学扰动关联。在中文版 InstructTTSEval 上,该方案将无需微调时的指令遵循率从34.5%提升至56.4%,有简单微调时从51.0%升至56.4%;受控重写使漂移率由40.4%降至15.4%。消融实验验证三机制互补性,且漂移分类体系可推广至非语音的指令生成任务。

原文摘要 · Abstract (English)

Instruct-TTS systems expand structured style labels into natural-language training instructions through LLM rewriting, yet we find that over 40% of unconstrained rewrites contain semantic drift that corrupts supervision and weakens generalization. We formalize this problem as instruction supervision instability and propose a data-centric stabilization recipe that jointly improves coverage and fidelity through three mechanisms: controllable instruction diversification for systematic expansion, LLM-based drift filtering for quality control, and attribute-aligned supervision that grounds prosody control in acoustic perturbations. On the Chinese split of InstructTTSEval, our recipe raises instruction-following from 34.5% without fine-tuning and 51.0% with naive fine-tuning to 56.4%, while constrained rewriting reduces drift from 40.4% to 15.4%. Ablations confirm the three mechanisms are complementary, and the drift taxonomy may generalize to instruction-driven generation beyond TTS.

语音合成指令监督语义漂移大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。