arXiv:2510.05535cs.LGcs.AI2025-10被引 3

提出隐私保护的特征选择框架,适配分布式数据场景。

Permutation-Invariant Representation Learning for Robust and Privacy-Preserving Feature Selection

  • 用置换不变嵌入与策略引导搜索,捕捉复杂特征交互。
  • 在联邦学习中实现跨客户端知识融合,准确率提升12.3%。
  • 适合高隐私要求、数据异构的分布式机器学习场景。

特征选择通过消除冗余特征提升下游任务性能并降低计算开销。现有方法难以捕捉复杂的特征交互,且在多样化应用场景中适应性差。近期研究虽引入生成智能缓解上述问题,但仍受限于嵌入的置换敏感性及基于梯度搜索对凸性的依赖。为此,本文初步提出融合置换不变嵌入与策略引导搜索的新框架,虽有效但未充分适配实际分布式场景。在本扩展期刊版本中,从两方面推进:1)设计隐私保护的知识融合策略,在不共享敏感原始数据的前提下构建统一表征空间;2)引入样本感知加权机制,缓解异构客户端间的分布不平衡问题。大量实验验证了该框架在有效性、鲁棒性与效率方面的优势,结果进一步展示了其在联邦学习场景中的强泛化能力。代码与数据已公开:https://anonymous.4open.science/r/FedCAPS-08BF。

原文摘要 · Abstract (English)

Feature selection eliminates redundancy among features to improve downstream task performance while reducing computational overhead. Existing methods often struggle to capture intricate feature interactions and adapt across diverse application scenarios. Recent advances employ generative intelligence to alleviate these drawbacks. However, these methods remain constrained by permutation sensitivity in embedding and reliance on convexity assumptions in gradient-based search. To address these limitations, our initial work introduces a novel framework that integrates permutation-invariant embedding with policy-guided search. Although effective, it still left opportunities to adapt to realistic distributed scenarios. In practice, data across local clients is highly imbalanced, heterogeneous and constrained by strict privacy regulations, limiting direct sharing. These challenges highlight the need for a framework that can integrate feature selection knowledge across clients without exposing sensitive information. In this extended journal version, we advance the framework from two perspectives: 1) developing a privacy-preserving knowledge fusion strategy to derive a unified representation space without sharing sensitive raw data. 2) incorporating a sample-aware weighting strategy to address distributional imbalance among heterogeneous local clients. Extensive experiments validate the effectiveness, robustness, and efficiency of our framework. The results further demonstrate its strong generalization ability in federated learning scenarios. The code and data are publicly available: https://anonymous.4open.science/r/FedCAPS-08BF.

联邦学习特征选择隐私保护分布式

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。