联邦个性化模型存在难检测的隐性失效,影响公平与对齐。
Silent Failures in Federated Personalization of Foundation Models
- 提出六类隐性失效模式,源于模型个性化与联邦约束的交互。
- 隐私保护训练无法确保可信部署,现有评估框架存在结构性盲区。
- 适合关注联邦学习安全、模型可信度的研究者与工程师。
基础模型通过联邦学习在分布式私有数据上进行个性化,正面临日益严格的上市后监控监管要求。我们指出,这种融合催生了一类未被充分认识的信任缺失问题,称为「隐性失效」,包括偏见放大、公平性崩溃和对齐退化,且由于联邦学习的隐私限制,这些失效难以被察觉。对现有基准的分析显示:联邦基准评估系统性能但难以揭示模型行为,而集中式信任基准虽能评估行为,却需要模型访问权限,与联邦隐私原则冲突。本文提出六种由基础模型个性化、数据分布偏移及联邦核心约束相互作用引发的隐性失效模式。分析表明,仅靠隐私保护训练不足以实现可信部署。最后提出隐私保护行为评估的研究议程,并建议将隐性失效作为可信联邦人工智能的标准诊断类别。
原文摘要 · Abstract (English)
Foundation models are increasingly personalized on decentralized private data through federated learning and are now deployed at scale under growing regulatory requirements for post-market monitoring. We argue that this convergence creates a distinct and under-recognized class of trustworthiness failures, which we term "Silent Failures." These include amplified bias, fairness collapse, and alignment erosion that may remain difficult to detect because federated learning's privacy constraints limit visibility into model behavior. A landscape analysis of existing benchmarks reveals a structural divide. Federated benchmarks evaluate system performance but provide limited insight into model behavior, whereas centralized trustworthiness benchmarks assess behavior but require model access incompatible with federated privacy. We introduce a taxonomy of six silent failure modes arising from the interaction of foundation model personalization, dataset shift, and core federated constraints. Our analysis shows that privacy-preserving training alone is insufficient for trustworthy deployment. We conclude with a research agenda for privacy-preserving behavioral evaluation and propose that silent failures become a standard diagnostic category for trustworthy federated artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。