arXiv:2606.00684eess.AScs.CL2026-06

通过分离表示中的关键成分,提升语音误读检测的准确性。

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection

论文配图:Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection
图 1 · 摘自论文原文
  • 构建子流框架,分离重要特征与上下文信息
  • 在真实数据集上实现零样本误读检测性能超越似然方法
  • 基于速度场设计几何诊断指标,适合语音异常检测场景

针对高维数据空间中目标观测的分布外检测问题,本文采用连续归一化流(CNFs)提出拉格朗日子流(LSF)框架,旨在分离并估计表示中相关成分的密度,同时将剩余成分作为上下文。通过语音合成模型实验发现,与其它深度生成模型类似,CNFs也存在“似然悖论”——对分布外样本错误地赋予高似然值。这源于生成模型的归纳偏置:更关注低层次结构细节而非高层次语义一致性。为缓解该现象,本文基于子流轨迹上的速度场设计了多种几何诊断信号,并据此构建用于零样本音素级误读检测的度量。在真实世界误读检测基准测试中,所提方法显著优于基于似然的方法。

原文摘要 · Abstract (English)

We address the problem of out-of-distribution (OOD) detection for target observations embedded in a subspace of the high dimensional data space. Using continuous normalizing flows (CNFs), we propose a Lagrangian sub-flow (LSF) framework designed to isolate and estimate the density for the relevant components in the representation and using the remaining components as context. Through experimentation with models for speech synthesis, we show that CNFs, similarly to other deep generative models (DGMs), are susceptible to the "likelihood paradox", where high likelihood is erroneously assigned to OOD samples. This is attributed to the inductive bias of DGMs that prioritize low-level structural details over high-level semantic coherence. To mitigate this phenomenon, we propose a number of geometric diagnostic signals based on the velocity field over the sub-flow trajectory. Based on these signals, we design metrics for the challenging task of zero-shot phoneme-level mispronunciation detection. Finally, we demonstrate the superiority of these metrics compared to likelihood-based methods on a real-world mispronunciation detection benchmark.

分布外检测生成模型语音分析几何诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。