用迁移学习提升小样本下伊辛模型的估计精度
Transfer Learning in High-dimensional Ising Models

- 先筛选有用辅助数据,再分两阶段估计避免负迁移
- 实测误差低于仅用目标数据或直接拼接数据的方法
- 适合小样本高维数据建模,尤其需借用外部数据时
在高维伊辛模型估计中,目标样本量常受限,如何有效利用未知相关性的辅助二值数据仍具挑战。为此,我们提出Trans-Ising方法,结合基于损失的源数据筛选规则与两阶段估计流程。首先通过目标数据的伪似然保留来识别有信息的辅助源,防止负迁移;随后通过合并节点的ℓ₁正则化逻辑回归获得初始估计,并使用折叠凸惩罚进行仅目标数据修正。理论上,我们建立了固定节点的ℓ₂与ℓ₁误差界、图选择的一致性以及筛选规则的条件一致性。通过大量模拟和真实数据分析,结果表明Trans-Ising在估计误差上优于仅使用目标数据或简单拼接数据的方法。
原文摘要 · Abstract (English)
In high-dimensional Ising model estimation, target sample sizes are often limited, and effectively using auxiliary binary datasets of unknown relevance remains challenging. To address this, we propose Trans-Ising, a transfer learning method that combines a loss-based source screening rule with a two-stage estimation procedure. The method first identifies informative auxiliary sources using held-out target pseudolikelihood to prevent negative transfer. It then computes an initial estimator via pooled nodewise $\ell_1$-regularized logistic regression, followed by a target-only correction step using a folded-concave penalty. Theoretically, we establish fixed-node $\ell_2$ and $\ell_1$ error bounds, exact graph selection consistency, and the conditional consistency of the screening rule. Through extensive simulations and real-data analyses, we demonstrate that Trans-Ising achieves lower estimation errors than both target-only estimation and naive data pooling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。