arXiv:2409.01410cs.LGstat.CO2024-09被引 3

为数据蒸馏建立任务导向的理论框架,解决其应用中的模糊性问题。

Dataset Distillation from First Principles: Integrating Core Information Extraction and Purposeful Learning

  • 基于目标任务定义数据蒸馏的优化问题,避免盲目选择方法。
  • 在医学数据分析中实现异构特征数据集的知识融合,提升小样本效果。
  • 帮助物理信息神经网络生成更符合物理规律的训练数据,缓解分布外误差。

数据蒸馏(DD)是一种日益重要的技术,旨在构建一个能捕捉训练数据核心信息的合成数据集,使在后者上训练的模型达到相当性能。尽管应用广泛,但其理论基础尚不成熟。现有方法多在通用基准上比较,缺乏针对具体学习任务的导向。本文提出一个形式化模型,强调必须明确与应用相关的推理任务,才能精确刻画底层优化问题。否则,数据蒸馏问题未被充分定义,算法选择仅凭经验。该形式化揭示了跨不同建模环境的新应用场景。我们通过这一新视角分析现有方法,指出其在准确性和最优操作忠实度上的优劣。最后,展示了两个现代场景下的数值结果:一是医疗数据分析中,合并具有交集但非完全相同的特征集的数据,以在小样本条件下构建更大数据集;二是针对物理信息神经网络(PINNs)在边界条件变化下的分布外误差,展示数据蒸馏提升物理一致性的潜力。本研究旨在建立一种新的研究范式,推动数据蒸馏的理论发展和新方法诞生。

原文摘要 · Abstract (English)

Dataset distillation (DD) is an increasingly important technique that focuses on constructing a synthetic dataset capable of capturing the core information in training data to achieve comparable performance in models trained on the latter. While DD has a wide range of applications, the theory supporting it is less well evolved. New methods of DD are compared on a common set of benchmarks, rather than oriented towards any particular learning task. In this work, we present a formal model of DD, arguing that a precise characterization of the underlying optimization problem must specify the inference task associated with the application of interest. Without this task-specific focus, the DD problem is under-specified, and the selection of a DD algorithm for a particular task is merely heuristic. Our formalization reveals novel applications of DD across different modeling environments. We analyze existing DD methods through this broader lens, highlighting their strengths and limitations in terms of accuracy and faithfulness to optimal DD operation. Finally, we present numerical results for two case studies important in contemporary settings. Firstly, we address a critical challenge in medical data analysis: merging the knowledge from different datasets composed of intersecting, but not identical, sets of features, in order to construct a larger dataset in what is usually a small sample setting. Secondly, we consider out-of-distribution error across boundary conditions for physics-informed neural networks (PINNs), showing the potential for DD to provide more physically faithful data. By establishing this general formulation of DD, we aim to establish a new research paradigm by which DD can be understood and from which new DD techniques can arise.

数据蒸馏小样本学习物理信息网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。