arXiv:2507.21205cs.LGcs.AI2025-07

让深度模型学会从真实世界杂乱数据中高效学习

Learning from Limited and Imperfect Data

  • 针对长尾分布数据设计抗模式崩溃的生成算法
  • 通过归纳正则化使少数类泛化能力接近主流类
  • 支持低标注甚至无标签下的高效领域自适应

现实世界数据分布(如互联网数据)与精心整理的数据集差异显著,常存在类别极度不均衡的问题。针对这一挑战,本文提出四类实用算法,使深度神经网络能在有限且不完美数据下有效学习。第一部分解决长尾数据下的生成模型训练问题,缓解模式崩溃,实现少数类图像的多样美学生成;第二部分通过归纳正则化策略,使尾部类别在不生成图像的前提下达到与头部类别相当的泛化能力;第三部分面向标注稀缺场景,优化相关评估指标的半监督学习;第四部分实现模型在极少量甚至无标签样本下的高效领域自适应。这些方法为真实世界数据应用提供了可扩展的鲁棒解决方案。

原文摘要 · Abstract (English)

The distribution of data in the world (eg, internet, etc.) significantly differs from the well-curated datasets and is often over-populated with samples from common categories. The algorithms designed for well-curated datasets perform suboptimally when used for learning from imperfect datasets with long-tailed imbalances and distribution shifts. To expand the use of deep models, it is essential to overcome the labor-intensive curation process by developing robust algorithms that can learn from diverse, real-world data distributions. Toward this goal, we develop practical algorithms for Deep Neural Networks which can learn from limited and imperfect data present in the real world. This thesis is divided into four segments, each covering a scenario of learning from limited or imperfect data. The first part of the thesis focuses on Learning Generative Models from Long-Tail Data, where we mitigate the mode-collapse and enable diverse aesthetic image generations for tail (minority) classes. In the second part, we enable effective generalization on tail classes through Inductive Regularization schemes, which allow tail classes to generalize as effectively as the head classes without requiring explicit generation of images. In the third part, we develop algorithms for Optimizing Relevant Metrics for learning from long-tailed data with limited annotation (semi-supervised), followed by the fourth part, which focuses on the Efficient Domain Adaptation of the model to various domains with very few to zero labeled samples.

长尾分布生成模型半监督学习领域自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。