arXiv:2608.19914cs.LG2026-08

用分布鲁棒优化融合异构数据,提升小样本下复杂网络重建精度。

Multi-Source Complex Network Reconstruction via Wasserstein Distributionally Robust Optimization and Algorithm Unrolling

论文配图:Multi-Source Complex Network Reconstruction via Wasserstein Distributionally Robust Optimization and Algorithm Unrolling
图 1 · 摘自论文原文
  • 基于加权沃斯泰因中位数融合多源数据,保留各来源几何结构。
  • 在小样本场景下显著优于7个基线方法,诊断性能提升明显。
  • 端到端可微架构自动调参,适合数据稀缺的神经影像等场景。

从数据中重构复杂网络拓扑是控制论与图信号处理的基础挑战,广泛应用于神经科学、传感器和社交网络。实际中目标域样本稀少而异构源域数据丰富,直接平均会因源间差异导致性能下降,使不同几何结构坍缩为有偏共识。本文提出多源沃斯泰因分布鲁棒优化(MS-WDRO)框架,利用沃斯泰因度量保持分布特性,通过加权沃斯泰因中位数构建几何合理的核心分布,并在其周围建立不确定性球体以应对残余不确定性。最小化最坏情况风险后,得到可高效求解的正则化拉普拉斯估计器,采用证明收敛的ADMM算法。理论方面,建立了经验中位数的有限样本集中界、朴素聚合的池化偏差下界,以及仅对源数量对数依赖的外样本超额风险界。通过将求解器展开为可微架构,实现超参数的端到端自适应校准,兼具可解释性与数据适应性。在合成基准和多中心ABIDE-I脑影像数据集上的实验表明,MS-WDRO在图恢复、样本效率及下游诊断能力上持续优于7个基线,在样本稀缺条件下提升最显著。

原文摘要 · Abstract (English)

Reconstructing complex network topologies from data is a fundamental challenge in cybernetics and graph signal processing, with applications in neuroscience, sensor, and social networks. In practice, target-domain samples are scarce while heterogeneous source-domain data are abundant. Fusing these sources is challenging: Euclidean averaging works for homogeneous sources but degrades sharply as inter-source divergence grows, collapsing distinct geometries into an inflated, biased consensus. We exploit the Wasserstein metric's distribution-preserving properties to counter heterogeneity while preserving each source's intrinsic geometry. We propose MS-WDRO, a multi-source Wasserstein distributionally robust graph learning framework that fuses heterogeneous sources via their weighted Wasserstein barycenter, a geometrically principled nominal distribution, then builds an ambiguity ball around it to hedge residual uncertainty. Minimizing worst-case risk yields a tractable regularized Laplacian estimator solved efficiently via a provably convergent ADMM scheme. We establish non-asymptotic guarantees: a finite-sample concentration bound for the empirical barycenter, a pooling bias lower bound proving naive aggregation is suboptimal, and an out-of-sample excess risk bound decaying at a parametric rate with only logarithmic dependence on source count. To calibrate hyperparameters governing robustness, sparsity, and source fusion, we unroll the solver into a differentiable architecture trained end-to-end, achieving data-adaptive calibration beyond cross-validation while retaining interpretability. Experiments on synthetic benchmarks and the multi-site ABIDE~I neuroimaging dataset show MS-WDRO consistently outperforms seven baselines in graph recovery, sample efficiency, and downstream diagnostic utility, with the largest gains in the sample-scarce regime.

网络重建分布鲁棒脑影像小样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。