通过分层投影实现自适应知识迁移,提升跨域学习的准确性和可解释性。
Hierarchical Projection for Adaptive Knowledge Transfer
- 构建分层贝叶斯先验,动态加权多个异构源数据
- 两阶段迁移:全局对齐+局部特征选择,减少负迁移
- 适用于高维生物医学数据,结果更稳定可解释
现代数据驱动应用越来越多地需要从多个异构来源中学习,目标数据集规模有限但存在相关领域信息。若盲目合并这些来源,当相关性不同时或存在虚假信号时,性能可能下降,这给可信的跨域学习带来根本挑战。我们提出投影迁移学习(ProjectionTL),一个统一框架,将分层贝叶斯建模与自适应投影相结合,实现选择性知识迁移。核心思想是分两级解耦迁移:首先,构建源引导的分层先验,利用数据驱动权重聚合各源信息,捕捉每个源与目标间的全局对齐;其次,通过后验投影步骤在特征层面精炼借用内容,仅保留与目标信号局部一致的坐标。该两阶段设计能同时完成源选择与特征选择,从而缓解负迁移并保持可解释性。ProjectionTL为跨域异构数据整合提供了原则性方法,连接统计建模与现代机器学习范式,实现稳健且可解释的迁移。通过模拟和真实生物医学应用验证,相比现有方法,其在准确性、稳定性与可解释性上均有提升。本框架在高维场景下具备可扩展性和通用性,为可信跨域学习提供策略。
原文摘要 · Abstract (English)
Modern data-driven applications increasingly involve learning from multiple heterogeneous sources, where a target dataset is limited but related information is available across domains. Naively combining these sources can degrade performance when relevance varies or spurious signals are present, posing a fundamental challenge for trustworthy cross-domain learning. We propose Projection Transfer Learning (ProjectionTL), a unified framework that integrates hierarchical Bayesian modeling with adaptive projection for selective knowledge transfer. The key idea is to decouple transfer at two levels: first, we construct a source-guided hierarchical prior that aggregates information across sources using data-driven weights, capturing global alignment between each source and the target; second, we refine this borrowing through a posterior-projection step that operates at the feature level, selectively retaining coordinates that exhibit local agreement with the target signal. This two-stage design enables the method to simultaneously perform source selection and feature selection, thereby mitigating negative transfer while preserving interpretability. ProjectionTL provides a principled approach to integrating heterogeneous data across domains, bridging statistical modeling and modern machine learning paradigms for robust and interpretable transfer. Through simulations and real-world biomedical applications, we demonstrate improved accuracy, stability, and interpretability compared to existing methods. Our framework offers a scalable and generalizable strategy for trustworthy cross-domain learning in high-dimensional settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。