让语音合成系统学会根据上下文推理说话方式,提升对话自然度。
ISCSLP 2026 CoT-TTS Challenge: Chain-of-Thought Reasoning for Context-Aware Text-to-Speech

- 通过链式思维推理,从上下文自动推断语调与风格
- 支持中英双语,提供大规模训练与评测数据集
- 适合做虚拟角色、有声书、影视配音等场景研究
近期文本到语音(TTS)技术显著提升了语音自然度、说话人相似性和可控性。然而,现有可控TTS系统大多依赖用户显式提供风格提示,在长而复杂的对话场景中难以自动判断应如何表达。本文提出ISCSLP 2026 CoT-TTS挑战,旨在评估系统是否能从上下文信息中推理出预期的说话方式,并生成与推理结果和语境一致的语音。挑战包含两个赛道:文本上下文感知的链式思维语音合成(text-context-aware CoT-TTS)与音频上下文感知的链式思维语音合成(audio-context-aware CoT-TTS)。我们基于语音丰富的媒体构建了大规模双语训练集,并提供经过精心筛选的评测数据用于排行榜比较。每个系统需输出链式思维推理分析与生成的语音波形。官方评估结合客观指标、多模态大模型评估及人工主观评分。为促进可复现性,我们提供基于0.6B Qwen3模型的推理代码与三阶段微调方案。该挑战有望推动上下文理解、链式思维推理与情感化语音生成在影视配音、有声读物、虚拟角色及语音对话代理等领域的研究。更多信息请访问:https://iscslp2026-cot-tts.github.io/challenge-website/
原文摘要 · Abstract (English)
Recent advances in text-to-speech (TTS) have greatly improved speech naturalness, speaker similarity, and controllability. However, most existing controllable TTS systems still rely on explicit user-provided style prompts, making it difficult to automatically determine how a sentence should be spoken in long and complex conversational scenarios. This proposal introduces the ISCSLP 2026 CoT-TTS Challenge, which aims to evaluate whether a system can infer the intended speaking manner from contextual information and generate speech consistent with both the reasoning output and the surrounding scene. The challenge contains two tracks: text-context-aware CoT-TTS and audio-context-aware CoT-TTS. We construct a large-scale bilingual training set from speech-rich media and provide carefully filtered evaluation data for leaderboard comparison. Each system is required to output both a chain-of-thought reasoning analysis and the generated speech waveform. The official evaluation combines objective metrics, multimodal LLM-based evaluation, and human subjective assessment. To facilitate reproducibility, we provide inference code together with a fine-tuning recipe for a 0.6B Qwen3-based model trained via a three-stage strategy. This challenge is expected to support research on context understanding, chain-of-thought reasoning, and expressive speech generation for applications such as film dubbing, audiobook production, virtual characters, and spoken dialogue agents. Further information about the associated challenge is available at:https://iscslp2026-cot-tts.github.io/challenge-website/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。