arXiv:2411.02921cs.LG2024-11

提出可证明泛化能力的动态分布学习框架,实现模型对数据流变化的自适应跟踪。

Theoretically Guaranteed Distribution Adaptable Learning

  • 通过编码特征边缘分布信息,突破最优传输限制,实现跨分布复用
  • 基于费希尔-劳距离分析整个分类器轨迹的泛化误差界
  • 适用于需要持续学习和可解释性的开放环境智能系统

在许多开放环境应用中,数据以流的形式持续采集,其分布随时间不断演变。如何设计能追踪这些动态分布且具备可证明泛化能力的算法,仍是重大挑战。为应对这一关键但极少被研究的问题,推动鲁棒人工智能发展,本文提出一种新型框架——分布可适配学习(Distribution Adaptable Learning, DAL)。该框架使模型能够有效追踪演化中的数据分布。通过编码特征边缘分布信息(EFMDI),突破了最优传输方法在刻画环境变化上的局限性,支持模型在多样分布间复用,显著提升DAL的可重用性和可演化性。为进一步增强模型可解释性,不仅分析了演化过程中局部步骤的泛化误差界,还基于费希尔-劳距离研究了整个分类器轨迹的泛化误差界。文中还给出了框架内的两个特例,附带优化方法与收敛性分析。在合成数据和真实世界数据分布演化任务上的实验结果验证了该框架的有效性与实用价值。

原文摘要 · Abstract (English)

In many open environment applications, data are collected in the form of a stream, which exhibits an evolving distribution over time. How to design algorithms to track these evolving data distributions with provable guarantees, particularly in terms of the generalization ability, remains a formidable challenge. To handle this crucial but rarely studied problem and take a further step toward robust artificial intelligence, we propose a novel framework called Distribution Adaptable Learning (DAL). It enables the model to effectively track the evolving data distributions. By Encoding Feature Marginal Distribution Information (EFMDI), we broke the limitations of optimal transport to characterize the environmental changes and enable model reuse across diverse data distributions. It can enhance the reusable and evolvable properties of DAL in accommodating evolving distributions. Furthermore, to obtain the model interpretability, we not only analyze the generalization error bound of the local step in the evolution process, but also investigate the generalization error bound associated with the entire classifier trajectory of the evolution based on the Fisher-Rao distance. For demonstration, we also present two special cases within the framework, together with their optimizations and convergence analyses. Experimental results over both synthetic and real-world data distribution evolving tasks validate the effectiveness and practical utility of the proposed framework.

持续学习分布演化泛化保证可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。