arXiv:2603.19798cs.SDcs.CL2026-03被引 2

让语音合成突破句子边界,支持多说话人、长文本和指令控制。

Borderless Long Speech Synthesis

  • 用分层标注结构实现从场景到音素的全流程控制
  • 结合思维链与维度丢弃,显著提升复杂指令遵循能力
  • 适合需要自然交互式语音生成的智能体应用

现有文本到语音(TTS)系统多为逐句合成或仅依赖纯文本对话,缺乏对全局上下文和副语言线索的理解,难以捕捉真实场景中的多说话人互动(如打断、重叠说话)、情绪变化轨迹及多样声学环境。我们提出面向代理的无边界长语音合成框架,作为统一能力集,覆盖VoiceDesigner、多说话人合成、指令语音合成和长文本合成。数据层面,采用“标注优先于筛选/清洗”的策略,设计名为Global-Sentence-Token的自上而下的多级标注体系;模型层面,采用连续分词器,并引入思维链(CoT)推理与维度丢弃,显著提升复杂条件下的指令遵循能力。系统天然具备代理特性:层级标注充当大语言模型代理与合成引擎间的结构化语义接口,构建从场景语义到音素细节的分层控制协议栈。文本因此成为信息完备、宽频带的控制通道,使前端大模型可将任意模态输入转化为结构化生成指令,推动范式从Text2Speech迈向无边界长语音合成。

原文摘要 · Abstract (English)

Most existing text-to-speech (TTS) systems either synthesize speech sentence by sentence and stitch the results together, or drive synthesis from plain-text dialogues alone. Both approaches leave models with little understanding of global context or paralinguistic cues, making it hard to capture real-world phenomena such as multi-speaker interactions (interruptions, overlapping speech), evolving emotional arcs, and varied acoustic environments. We introduce the Borderless Long Speech Synthesis framework for agent-centric, borderless long audio synthesis. Rather than targeting a single narrow task, the system is designed as a unified capability set spanning VoiceDesigner, multi-speaker synthesis, Instruct TTS, and long-form text synthesis. On the data side, we propose a "Labeling over filtering/cleaning" strategy and design a top-down, multi-level annotation schema we call Global-Sentence-Token. On the model side, we adopt a backbone with a continuous tokenizer and add Chain-of-Thought (CoT) reasoning together with Dimension Dropout, both of which markedly improve instruction following under complex conditions. We further show that the system is Native Agentic by design: the hierarchical annotation doubles as a Structured Semantic Interface between the LLM Agent and the synthesis engine, creating a layered control protocol stack that spans from scene semantics down to phonetic detail. Text thereby becomes an information-complete, wide-band control channel, enabling a front-end LLM to convert inputs of any modality into structured generation commands, extending the paradigm from Text2Speech to borderless long speech synthesis.

语音合成长语音智能体指令控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。