arXiv:2608.02673cs.SDcs.AI2026-08

用标签精确控制语音编辑,支持文本、情绪、语调和停顿修改。

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

论文配图:dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
图 1 · 摘自论文原文
  • 通过带标签的文本指令精准定位编辑区域,避免时间对齐。
  • 在四个语音控制维度上实现领先指令遵循率与局部保留效果。
  • 适合需要精细语音修改的创作人员和语音合成研究者。

语音编辑在内容创作中需要对编辑内容和位置都有精确控制。自由形式的自然语言虽灵活但易模糊,难以明确操作目标。本文提出一种基于转录文本的结构化编辑指令,使用类似XML的标签显式指定操作类型并定位到文本片段或边界,构建可外部检视的语义时间线,无需显式时间对齐。我们基于连续自回归模型dots.tts构建了dots.tts.edit编辑器,支持四类控制:词汇内容、情感表达、语调与语速、时间分段,分别通过文本、情绪、韵律和停顿编辑实现。任务特定的数据流水线构建操作与范围可控的成对样本,同时保留目标区域外的源上下文。我们还引入doteBench,一个双语评估套件,衡量指令遵循度、局部保留性和音频质量,涵盖四项控制及其组合。实验显示,在五类编辑任务中整体指令遵循率和局部保留性能领先,音频质量与现有开源系统相当。在三个Seed-TTS-Eval数据集上,零样本语音识别误差率和说话人相似性与基线模型无显著差异。

原文摘要 · Abstract (English)

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, parameters, or target region underspecified. We study a precise and explicit interface for speech editing: a transcript-grounded structural edit instruction with XML-style tags explicitly specifies typed operations and localizes them to transcript spans or boundaries. This semantic timeline avoids explicit timestamp alignment and provides an externally inspectable contract for compositional edits. We instantiate the interface in dots$.$tts$.$edit, an editor adapted from the continuous autoregressive dots$.$tts foundation model. Four representative speech-creation controls cover lexical content, affective expression, pitch and speaking-rate delivery, and temporal phrasing through text, emotion, prosody, and pause editing. Task-specific data pipelines construct operation- and scope-controlled pairs while retaining source-derived context outside each target region. We further introduce doteBench, a bilingual evaluation suite that measures precise instruction following, local preservation, and audio quality across the four controls and their composition. Experiments show leading overall instruction following and local preservation across its five editing categories, while audio quality remains comparable to existing open-source systems. Across three Seed-TTS-Eval shards, the model shows negligible differences from the base model in zero-shot TTS recognition error rate and speaker similarity.

语音编辑自回归模型语音合成指令控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。