arXiv:2605.14301cs.LGstat.ML2026-05

用文本描述构建领域先验,解决小样本场景下的领域自适应难题

Language-Induced Priors for Domain Adaptation

论文配图:Language-Induced Priors for Domain Adaptation
图 1 · 摘自论文原文
  • 基于大模型生成目标域文本先验,指导源域选择
  • 在数据稀缺时提升预测性能,冷启动误差接近理想情况
  • 适用于各类参数化模型,适合小样本、无标签场景

领域自适应在冷启动场景下面临根本性矛盾:当目标数据稀少时,统计方法难以区分相关与无关源域,常导致负迁移。本文利用常被忽略的目标域专家文本描述,提出一种概率框架,将语义描述转化为语言诱导先验(LIP),通过预训练大模型学习偏好。LIP被集成进期望最大化算法,用于识别源域相关性。该方法适用于任意具备似然函数的参数化模型,能在目标信号弱时引导源域选择,并随样本积累逐步优化。理论上证明,若先验正确,估计器误差近似于理想冷启动均方误差,且无论先验质量如何,均保持渐近一致性。实验验证了其在高斯估计(描述性任务)、C-MAPSS(预测性任务)和MuJoCo Hopper(决策性任务)上的有效性。

原文摘要 · Abstract (English)

Domain adaptation faces a fundamental paradox in the cold-start regime. When target data is scarce, statistical methods fail to distinguish relevant source domains from irrelevant ones, which often leads to negative transfer. In this paper, we address this challenge by leveraging expert textual descriptions of the target domain, a resource that is often available but overlooked. We propose a probabilistic framework that translates these semantic descriptions into a choice model, namely a Language-Induced Prior (LIP), that learns the preferences from a pretrained Large Language Model (LLM). The LIP is then integrated into an Expectation-Maximization algorithm to identify source relevance. Methodologically, this framework is compatible with any parametric model where a likelihood is available. It allows the LIP to guide the selection of sources when target signals are weak, while gradually refining these choices as samples accumulate. Theoretically, we prove that the estimator roughly matches an oracle cold-start MSE under a correct prior, while remaining asymptotically consistent regardless of the quality of the LIP. Empirically, we validated the framework on a descriptive (Gaussian estimation), a predictive (C-MAPSS dataset), and a prescriptive task (MuJoCo hopper).

领域自适应语言先验冷启动大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。