用相似度+大模型检测语音合成中的说话人漂移问题
A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech
- 通过重叠段落的余弦相似度判断说话人一致性
- 在多个SOTA大模型上验证了检测有效性
- 首个自动检测框架,适合语音合成质量评估者
基于扩散的文本到语音(TTS)模型虽具备高自然度与表现力,但常出现说话人漂移——单个语句中感知说话人身份的细微渐变。这一未被充分研究的现象损害了长语音或交互场景下的语音连贯性。本文提出首个自动检测说话人漂移的框架,将问题建模为语句级说话人一致性的二分类任务。方法计算合成语音重叠段落间的余弦相似度,并利用大语言模型(LLMs)对结构化表征进行漂移评估。理论证明余弦检测的有效性,且发现说话人嵌入在单位球面上存在有意义的几何聚类。为支持评估,构建了高质量合成基准数据集,包含人工验证的说话人漂移标注。多款先进大模型的实验验证了该嵌入-推理流程的可行性。本工作确立说话人漂移为独立研究课题,连接几何信号分析与大模型感知推理,推动现代TTS发展。
原文摘要 · Abstract (English)
Recent diffusion-based text-to-speech (TTS) models achieve high naturalness and expressiveness, yet often suffer from speaker drift, a subtle, gradual shift in perceived speaker identity within a single utterance. This underexplored phenomenon undermines the coherence of synthetic speech, especially in long-form or interactive settings. We introduce the first automatic framework for detecting speaker drift by formulating it as a binary classification task over utterance-level speaker consistency. Our method computes cosine similarity across overlapping segments of synthesized speech and prompts large language models (LLMs) with structured representations to assess drift. We provide theoretical guarantees for cosine-based drift detection and demonstrate that speaker embeddings exhibit meaningful geometric clustering on the unit sphere. To support evaluation, we construct a high-quality synthetic benchmark with human-validated speaker drift annotations. Experiments with multiple state-of-the-art LLMs confirm the viability of this embedding-to-reasoning pipeline. Our work establishes speaker drift as a standalone research problem and bridges geometric signal analysis with LLM-based perceptual reasoning in modern TTS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。