arXiv:2605.01907stat.MLcs.LG2026-05中稿 · ICML

提出新方法精准识别任务聚类并提升多任务学习效果。

Adaptive Estimation and Inference in Semi-parametric Heterogeneous Clustered Multitask Learning via Neyman Orthogonality

论文配图:Adaptive Estimation and Inference in Semi-parametric Heterogeneous Clustered Multitask Learning via Neyman Orthogonality
图 1 · 摘自论文原文
  • 用正交损失与自适应融合惩罚,动态校准聚类结构。
  • 高概率精确恢复聚类,收敛速度与簇大小成正比。
  • 适合有异质性、复杂结构的多任务学习问题。

研究半参数设定下的聚类多任务学习,其中任务共享目标参数的潜在聚类结构,但具有异质的、可能无限维的干扰成分。现有方法通常依赖对齐特征空间或同质任务结构,难以应对此类异质性。本文提出一种自适应融合正交估计器,结合奈曼正交损失与数据驱动的成对融合惩罚。通过任务特异性预估计校准融合惩罚,并将自适应聚合与正交化相结合,有效缓解干扰参数估计误差的影响。理论上,所提估计器以高概率实现潜在聚类的精确恢复,并达到与簇大小成比例的联合参数收敛速率。同时证明其渐近正态性,且渐近性能等价于已知真实聚类的预言者程序。实验表明,在多种模拟设置中持续优于强基线;在美国家庭用电消费的真实案例中,成功揭示有意义的区域电价弹性聚类,验证了方法的有效性。

原文摘要 · Abstract (English)

We study clustered multitask learning in a semiparametric setting where tasks share a latent cluster structure in their target parameters but exhibit heterogeneous, potentially infinite-dimensional nuisance components. Such heterogeneity poses a major challenge for existing multitask learning methods, which typically rely on aligned feature spaces or homogeneous task structures. To address this challenge, we propose an adaptive fused orthogonal estimator that integrates Neyman-orthogonal losses with data-driven pairwise fusion penalties. Our framework leverages task-specific pilot estimates to calibrate the fusion penalties and combines adaptive aggregation with orthogonalization to mitigate the impact of nuisance-parameter estimation error. Theoretically, we show that the proposed estimator achieves exact recovery of the latent clustering with high probability and attains pooled parametric convergence rates proportional to cluster size. Moreover, we establish asymptotic normality and show that, asymptotically, our estimator matches the performance of an oracle procedure that knows the true clustering in advance. Empirically, we show that the proposed method consistently outperforms strong baselines in various simulation setups. A real-world application to U.S. residential energy consumption demonstrates the effectiveness of our approach in uncovering meaningful regional clustering in electricity price elasticity, showcasing the efficacy of our method.

多任务学习聚类识别正交估计半参数模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。