arXiv:2504.04017cs.LG2025-04综述被引 2

跨领域少样本学习面临挑战与机遇,助力数据稀缺场景建模。

A Comprehensive Survey of Challenges and Opportunities of Few-Shot Learning Across Multiple Domains

  • 分析音频、图像、文本等多领域少样本学习的差异与共性。
  • 指出少样本学习在疾病爆发等紧急场景中可快速构建诊断模型。
  • 为不同应用场景选择合适方法提供理论依据,适合研究者参考。

在新领域不断涌现、机器学习持续应用于新任务的背景下,模型训练面临样本数量不足的挑战。传统机器学习依赖大量数据,但在新发传染病(如新冠疫情)初期,既难以获取足够样本,也缺乏对疾病的深入理解,导致无法快速构建有效的训练数据。少样本学习为此提供可行方案。本文系统梳理了音频、图像、文本及其组合等主要领域中少样本学习面临的挑战与机遇,分析其优势与局限。深入理解这些差异有助于针对不同领域和应用场景选择合适的少样本方法,推动其在医疗诊断等紧急情况下的快速应用。

原文摘要 · Abstract (English)

In a world where new domains are constantly discovered and machine learning (ML) is applied to automate new tasks every day, challenges arise with the number of samples available to train ML models. While the traditional ML training relies heavily on data volume, finding a large dataset with a lot of usable samples is not always easy, and often the process takes time. For instance, when a new human transmissible disease such as COVID-19 breaks out and there is an immediate surge for rapid diagnosis, followed by rapid isolation of infected individuals from healthy ones to contain the spread, there is an immediate need to create tools/automation using machine learning models. At the early stage of an outbreak, it is not only difficult to obtain a lot of samples, but also difficult to understand the details about the disease, to process the data needed to train a traditional ML model. A solution for this can be a few-shot learning approach. This paper presents challenges and opportunities of few-shot approaches that vary across major domains, i.e., audio, image, text, and their combinations, with their strengths and weaknesses. This detailed understanding can help to adopt appropriate approaches applicable to different domains and applications.

少样本学习跨领域医学应用综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。