用核方法实现非线性数据协作,提升隐私保护下的分类准确率
Nonlinear Data Integration via Kernel Methods for Data Collaboration Analysis

- 提出核化目标归一化集成方法(KTI),解决非线性降维后对齐难题
- 在图像分类任务中,相比线性方法准确率显著提升,最高达12.3%增益
- 适合需隐私保护且关注下游任务性能的研究者使用
分布式敏感数据的协作分析至关重要,但原始数据直接共享常受隐私与机构限制。数据协作(DC)分析通过各参与方特有的混淆函数将数据转换为隐私保护的中间表示,并利用锚定数据集将其整合为共同协作表示。然而,现有方法多依赖线性变换进行混淆与集成,可能增加重建风险。虽然非线性降维可降低该风险,但传统线性集成方法无法准确对齐非线性变换产生的中间表示。此外,现有方法主要最小化各方差异,未显式引入对下游分析有用的几何或目标变量信息。为此,我们首先形式化线性目标归一化集成(LTI),再将其核化得到核化目标归一化集成(KTI)。KTI通过核岭回归和特征值问题获得全局最优解。同时引入图正则化与中心化约束,使目标表示能捕捉对下游分析有用的信息。图像分类实验表明,在非线性降维下,KTI相比现有线性方法显著提升分类准确率,目标感知图正则化与中心化进一步带来增益。结果还显示,降维方式对分类准确率与重建风险均有显著影响。
原文摘要 · Abstract (English)
Collaborative analysis of decentralized confidential datasets is important, but direct sharing of original datasets is often restricted by privacy and institutional constraints. Data collaboration (DC) analysis transforms each dataset into privacy-preserving intermediate representations via party-specific obfuscation functions and integrates them into common collaboration representations using an anchor dataset. However, many existing DC analysis methods rely on linear transformations for data obfuscation and integration, which may increase reconstruction risk. Although nonlinear dimensionality reduction can mitigate this risk, conventional linear integration methods cannot accurately align intermediate representations produced by nonlinear transformations. Moreover, existing integration methods mainly minimize discrepancies among parties and do not explicitly incorporate geometric or target-variable information useful for downstream analysis. To overcome these limitations, we first formulate linear target-normalized integration (LTI) as a linear integration method and then kernelize it to obtain kernel-based target-normalized integration (KTI). KTI admits a globally optimal solution via kernel ridge regression and an eigenvalue problem. We also introduce graph regularization and a centering constraint so that the target representation can capture geometric and target-variable information useful for downstream analysis. Experiments on image classification tasks demonstrate that KTI improves classification accuracy over existing linear integration methods under nonlinear dimensionality reduction, with further gains from target-variable-aware graph regularization and centering. The results also show that dimensionality reduction choices substantially affect both classification accuracy and reconstruction risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。