arXiv:2509.16010cs.SDcs.AI2025-09中稿 · ICASSP 2026

用轻量级适配实现私密语音克隆,兼顾个性化与通信效率

Fed-PISA: Federated Voice Cloning via Personalized Identity-Style Adaptation

  • 分离音色与风格,仅上传轻量风格参数,降低通信开销
  • 通过相似风格用户协作,提升语音表达力和说话人相似度
  • 适合隐私敏感的语音生成场景,如智能客服、语音助手

文本到语音(TTS)中的语音克隆旨在用少量目标说话人数据生成富有表现力且个性化的语音。联邦学习(FL)为该任务提供了协作且保护隐私的框架,但现有方法存在通信成本高、抑制风格异质性的问题,导致个性化不足。为此,我们提出Fed-PISA,即联邦个性化音色-风格适配。为降低通信成本,引入解耦的低秩适配(LoRA)机制:说话人音色在本地通过私有ID-LoRA保留,仅上传轻量级风格-LoRA至服务器,从而最小化参数传输。为利用异质性,设计受协同过滤启发的聚合方法,通过学习风格相似的同伴数据,为每个客户端生成定制化模型。实验表明,Fed-PISA在风格表现力、自然度和说话人相似度上均优于标准联邦基线,且通信开销极低。

原文摘要 · Abstract (English)

Voice cloning for Text-to-Speech (TTS) aims to generate expressive and personalized speech from text using limited data from a target speaker. Federated Learning (FL) offers a collaborative and privacy-preserving framework for this task, but existing approaches suffer from high communication costs and tend to suppress stylistic heterogeneity, resulting in insufficient personalization. To address these issues, we propose Fed-PISA, which stands for Federated Personalized Identity-Style Adaptation. To minimize communication costs, Fed-PISA introduces a disentangled Low-Rank Adaptation (LoRA) mechanism: the speaker's timbre is retained locally through a private ID-LoRA, while only a lightweight style-LoRA is transmitted to the server, thereby minimizing parameter exchange. To harness heterogeneity, our aggregation method, inspired by collaborative filtering, is introduced to create custom models for each client by learning from stylistically similar peers. Experiments show that Fed-PISA improves style expressivity, naturalness, and speaker similarity, outperforming standard federated baselines with minimal communication costs.

语音克隆联邦学习低秩适配隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。