将语音生成源重新定义为多因素组合,提升新组合识别能力。
Open-Set Source Tracing as Compositional Factors via Structured Prototypes

- 用结构化正交原型降低类别重叠与类内差异
- 通过子空间划分实现架构、数据等因子解耦
- 在少样本开放集场景下显著优于传统方法
现有研究将语音生成源简单等同于生成架构,但本文提出将源重新定义为架构、训练数据及其他影响生成语音的训练因素的组合。为此,提出基于结构化正交原型的框架,以最小化类别间重叠和类内方差。采用子空间划分策略将嵌入分解为架构子空间、数据子空间,剩余子空间捕捉随机性变化,实现对未见因子组合的“组合泛化”。该方法在部分可见源上表现更优,并在完全开放集场景中保持鲁棒性。在MLAAD上的少样本开放集识别测试显示,本方法显著优于角度间距基线。
原文摘要 · Abstract (English)
Recent research expands beyond binary anti-spoofing with the emergence of Source Tracing, the task of identifying the specific generative origins of synthetic speech. However, current research often equates a "source" with its generative architecture. We propose redefining a source as a compositional tuple of Architecture, Training Data, and other training factors affecting the generated speech. We propose a framework using Structured Orthonormal Prototypes to minimize class overlap and intra-class variance. Our Subspace Partitioning strategy splits the embedding into architecture and data subspaces, while a residual subspace captures stochastic variability, enabling "compositional generalization" for novel factor combinations. This approach improves performance for partially seen sources and maintains robustness in fully open-set scenarios. MLAAD evaluations for Few-Shot open-set Identification show our approach significantly outperforms angular-margin baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。