arXiv:2510.18680cs.LG2025-10NeurIPS被引 4

用投票机制让模型学会通用表示,不依赖具体任务标签。

Learning Task-Agnostic Representations through Multi-Teacher Distillation

  • 设计投票式目标函数,自动融合多教师多样性特征。
  • 在文本、视觉和分子建模中提升下游任务性能。
  • 无需任务标签,适合跨领域通用表示学习。

将复杂输入转化为可处理的表示是多个领域的关键步骤。不同架构、损失函数、输入模态和数据集催生出多种嵌入模型,各自捕捉输入的不同特性。多教师蒸馏利用这种多样性丰富表示,但通常仍针对特定任务。本文提出基于“多数投票”目标函数的任务无关框架。我们证明该函数受学生与教师嵌入间互信息的约束,从而构建出不依赖任务标签或先验知识的任务无关蒸馏损失。在文本、视觉模型及分子建模中的评估表明,该方法有效利用教师多样性,生成的表示能显著提升分类、聚类、回归等下游任务表现。此外,我们训练并发布了当前最优的嵌入模型,在多种模态下均增强下游性能。

原文摘要 · Abstract (English)

Casting complex inputs into tractable representations is a critical step across various fields. Diverse embedding models emerge from differences in architectures, loss functions, input modalities and datasets, each capturing unique aspects of the input. Multi-teacher distillation leverages this diversity to enrich representations but often remains tailored to specific tasks. In this paper, we introduce a task-agnostic framework based on a ``majority vote" objective function. We demonstrate that this function is bounded by the mutual information between student and teachers' embeddings, leading to a task-agnostic distillation loss that eliminates dependence on task-specific labels or prior knowledge. Our evaluations across text, vision models, and molecular modeling show that our method effectively leverages teacher diversity, resulting in representations enabling better performance for a wide range of downstream tasks such as classification, clustering, or regression. Additionally, we train and release state-of-the-art embedding models, enhancing downstream performance in various modalities.

表示学习多教师蒸馏任务无关通用嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。