通过几何学习定位跨语言合成语音的情绪与操纵源。
Towards Attribution of Generators and Emotional Manipulation in Cross-Lingual Synthetic Speech using Geometric Learning
- 融合语音基础模型与声学特征,用曲率自适应投影实现多任务追踪。
- 在中英文合成语音数据集上,情绪与操纵源识别准确率显著提升。
- 首个针对合成语音多属性追踪的曲率自适应框架,适合安全检测研究者。
本文针对合成语音中情绪与操纵痕迹的细粒度溯源问题展开研究。我们提出假设:结合语音基础模型(SFMs)捕捉的语义-韵律线索与听觉表示中的细粒度频谱动态,可实现更精准的溯源。为此,我们设计了MiCuNet——一种多任务框架,通过混合曲率投影机制(跨越双曲、欧氏、球面空间),在可学习的时间门控引导下融合SFMs嵌入与谱图特征。该方法在包含中英文子集的EmoFake数据集(EFD)上,同时预测原始情绪、操控后情绪及操纵来源。实验表明,MiCuNet在各类指标上均优于传统融合策略,且首次实现针对合成语音多属性追踪的曲率自适应建模,为虚假语音溯源提供了新范式。
原文摘要 · Abstract (English)
In this work, we address the problem of finegrained traceback of emotional and manipulation characteristics from synthetically manipulated speech. We hypothesize that combining semantic-prosodic cues captured by Speech Foundation Models (SFMs) with fine-grained spectral dynamics from auditory representations can enable more precise tracing of both emotion and manipulation source. To validate this hypothesis, we introduce MiCuNet, a novel multitask framework for fine-grained tracing of emotional and manipulation attributes in synthetically generated speech. Our approach integrates SFM embeddings with spectrogram-based auditory features through a mixed-curvature projection mechanism that spans Hyperbolic, Euclidean, and Spherical spaces guided by a learnable temporal gating mechanism. Our proposed method adopts a multitask learning setup to simultaneously predict original emotions, manipulated emotions, and manipulation sources on the EmoFake dataset (EFD) across both English and Chinese subsets. MiCuNet yields consistent improvements, consistently surpassing conventional fusion strategies. To the best of our knowledge, this work presents the first study to explore a curvature-adaptive framework specifically tailored for multitask tracking in synthetic speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。