解释零样本预测为何能泛化,揭示其核心机制
A Generalization Theory for Zero-Shot Prediction
- 构建理论框架分析零样本预测的泛化能力
- 识别出模型学习的目标变量与条件独立关系
- 适合研究自监督学习与泛化理论的学者
机器学习与人工智能中的现代泛化范式包括预训练一个任务无关的基础模型,通常通过自监督和多模态对比学习获得。这些表示可用于在无标签数据的下游任务上进行预测。本文提出一个理论框架,以更好地理解这一方法——零样本预测。我们识别出零样本预测旨在学习或顺带学习的目标量,以及使其具备泛化能力的关键条件独立关系。
原文摘要 · Abstract (English)
A modern paradigm for generalization in machine learning and AI consists of pre-training a task-agnostic foundation model, generally obtained using self-supervised and multimodal contrastive learning. The resulting representations can be used for prediction on a downstream task for which no labeled data is available. We present a theoretical framework to better understand this approach, called zero-shot prediction. We identify the target quantities that zero-shot prediction aims to learn, or learns in passing, and the key conditional independence relationships that enable its generalization ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。