保护隐私的个性化虚拟人生成,用轻量适配器实现分布式训练。
PrivFedTalk: Privacy-Aware Federated Diffusion with Identity-Stable Adapters for Personalized Talking-Head Generation
- 共享扩散模型+本地轻量适配器,不传原始数据。
- 多客户端训练下身份稳定,画面无抖动、不走形。
- 适合低算力设备,支持隐私保护的联合建模。
基于扩散模型的虚拟人说话头生成技术发展迅速,但通常依赖集中式人脸视频与语音数据集,引发严重隐私问题。个性化生成中,身份相关数据尤为敏感,难以跨用户或设备共享。本文提出隐私感知的联邦生成框架 PrivFedTalk,结合条件潜空间扩散模型与参数高效的标识适配机制。在客户端间共享扩散主干网络,各客户端利用本地私有音视频数据训练轻量级 LoRA 标识适配器,避免原始数据暴露并降低通信开销。针对客户端分布异构问题,提出身份稳定联邦聚合(ISFA),基于设备端的身份一致性与时间稳定性估计,计算隐私安全的可靠性权重以加权更新。引入时序去噪一致性(TDC)正则化,有效减少联邦去噪过程中的帧间漂移、闪烁与身份偏移。为降低更新侧隐私风险,对适配器更新应用安全聚合与客户端级差分隐私。系统支持低显存 GPU 执行及异构共享硬件上的多 GPU 客户端并行训练。在多种训练与聚合条件下,与 FedAvg、FedProx 对比实验表明,该框架实现稳定联邦优化,可在资源受限环境下完成端到端训练与评估。结果验证了隐私感知个性化虚拟人生成在联邦环境中的可行性,同时指出未来需建立更标准的组件级、隐私-效用与定性评估体系。
原文摘要 · Abstract (English)
Talking-head generation has advanced rapidly with diffusion-based generative models, but training usually depends on centralized face-video and speech datasets, raising major privacy concerns. The problem is more acute for personalized talking-head generation, where identity-specific data are highly sensitive and often cannot be pooled across users or devices. PrivFedTalk is presented as a privacy-aware federated framework for personalized talking-head generation that combines conditional latent diffusion with parameter-efficient identity adaptation. A shared diffusion backbone is trained across clients, while each client learns lightweight LoRA identity adapters from local private audio-visual data, avoiding raw data sharing and reducing communication cost. To address heterogeneous client distributions, Identity-Stable Federated Aggregation (ISFA) weights client updates using privacy-safe scalar reliability signals computed from on-device identity consistency and temporal stability estimates. Temporal-Denoising Consistency (TDC) regularization is introduced to reduce inter-frame drift, flicker, and identity drift during federated denoising. To limit update-side privacy risk, secure aggregation and client-level differential privacy are applied to adapter updates. The implementation supports both low-memory GPU execution and multi-GPU client-parallel training on heterogeneous shared hardware. Comparative experiments on the present setup across multiple training and aggregation conditions with PrivFedTalk, FedAvg, and FedProx show stable federated optimization and successful end-to-end training and evaluation under constrained resources. The results support the feasibility of privacy-aware personalized talking-head training in federated environments, while suggesting that stronger component-wise, privacy-utility, and qualitative claims need further standardized evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。