arXiv:2512.15759cs.LG2025-12

通过语义约束提升联邦学习收敛速度与隐私-效用平衡,实测提速22%。

Semantic-Constrained Federated Aggregation: Convergence Theory and Privacy-Utility Bounds for Knowledge-Enhanced Distributed Learning

  • 引入语义约束机制,用领域知识指导模型聚合,避免无效更新。
  • 在制造数据上实现22%更快收敛,模型偏差降低41.3%,约束违规率<0.05时性能保持90%。
  • 理论证明收敛率,给出隐私-效用权衡边界,适合工业级分布式学习场景。

联邦学习在非独立同分布数据下收敛缓慢。现有方法对客户端更新一视同仁,忽视语义有效性。本文提出语义约束联邦聚合(SCFA),将领域知识约束融入分布式优化,首次建立基于约束的联邦学习收敛理论,证明收敛速率可达O(1/sqrt(T) + rho),其中rho为约束违反率。分析显示,约束使有效数据异质性降低41%,通过假设空间压缩因子theta=0.37改善隐私-效用平衡。在(epsilon,delta)-差分隐私下,epsilon=10时,约束正则化使性能损失仅3.7%,优于标准联邦学习的12.1%损失,提升2.7倍。在包含118万样本、968个传感器特征的Bosch生产数据上验证,构建基于ISA-95和MASON本体的知识图谱,包含3000条约束。实验表明,收敛速度提升22%,模型分歧减少41.3%;当rho<0.05时维持90%最优性能,rho>0.18则导致灾难性失效。理论预测与实证结果高度一致,各关系R²>0.90。

原文摘要 · Abstract (English)

Federated learning enables collaborative model training across distributed data sources but suffers from slow convergence under non-IID data conditions. Existing solutions employ algorithmic modifications treating all client updates identically, ignoring semantic validity. We introduce Semantic-Constrained Federated Aggregation (SCFA), a theoretically-grounded framework incorporating domain knowledge constraints into distributed optimization. We prove SCFA achieves convergence rate O(1/sqrt(T) + rho) where rho represents constraint violation rate, establishing the first convergence theory for constraint-based federated learning. Our analysis shows constraints reduce effective data heterogeneity by 41% and improve privacy-utility tradeoffs through hypothesis space reduction by factor theta=0.37. Under (epsilon,delta)-differential privacy with epsilon=10, constraint regularization maintains utility within 3.7% of non-private baseline versus 12.1% degradation for standard federated learning, representing 2.7x improvement. We validate our framework on manufacturing predictive maintenance using Bosch production data with 1.18 million samples and 968 sensor features, constructing knowledge graphs encoding 3,000 constraints from ISA-95 and MASON ontologies. Experiments demonstrate 22% faster convergence, 41.3% model divergence reduction, and constraint violation thresholds where rho<0.05 maintains 90% optimal performance while rho>0.18 causes catastrophic failure. Our theoretical predictions match empirical observations with R^2>0.90 across convergence, privacy, and violation-performance relationships.

联邦学习语义约束隐私保护工业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。