arXiv:2508.08644cs.CV2025-08被引 2

用熵最小化提升视觉语言模型在少样本下的鲁棒性

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation

  • 构建共享流形并最小化其熵,增强跨模态特征一致性
  • 在低数据场景下显著提升下游任务性能,泛化误差更小
  • 无需修改主干网络,可直接嵌入多种蒸馏框架

知识蒸馏是知识迁移的成熟技术,近年来在大规模视觉语言模型(VLMs)背景下重新受到关注。然而,视觉语言知识蒸馏通常需要充足训练数据以实现对模糊或边界邻近样本的鲁棒泛化,这些样本具有高预测不确定性。但在实际场景中,获取大规模、特定任务的数据往往不切实际。为解决不确定性与跨模态特征表示纠缠带来的挑战,本文提出对齐流形熵(AME),旨在实现真实条件下的鲁棒泛化。AME 在重构的共享流形上应用熵最小化,通过一对投影函数连接多模态数据(图像与文本),促进跨模态特征表示的结构压缩。该方法可在低数据条件下实现鲁棒知识蒸馏,且无需修改主干网络结构,可作为即插即用模块兼容多种视觉语言蒸馏框架。理论分析表明,将知识蒸馏与共享流形上的熵最小化结合,可获得更紧的泛化误差界。大量实验验证,无论在何种蒸馏架构与训练设置下,AME 均能持续提升知识蒸馏的鲁棒性,在广泛下游任务中表现更优。

原文摘要 · Abstract (English)

Knowledge distillation is a long-established technique for knowledge transfer, and has regained attention in the context of the recent emergence of large vision-language models (VLMs). However, vision-language knowledge distillation often requires sufficient training data to achieve robust generalization on amples with ambiguous or boundary-adjacent representations, which are associated with high predictive uncertainty. Critically, collecting such large-scale, task-specific data for training is often impractical in real-world scenarios. To address this major challenge arising from the entanglement of uncertainty and cross-modal feature representation, we propose Aligned Manifold Entropy for Robust Vision-Language Distillation (AME), aiming to achieve robust generalization under real-world conditions. AME applies entropy minimization over a reconfigured shared manifold, where multi-modal data (i.e., image and text) are bridged through a pair of projection functions, conducive to structural compression for cross-modal feature representations. This enables robust knowledge distillation under low-data regimes, while requiring no architectural modifications to the backbone. As a result, it can serve as a plug-and-play module compatible with a wide range of vision-language distillation frameworks. Notably, our theoretical analysis reveals that integrating knowledge distillation with entropy minimization over the shared manifold leads to a tighter generalization error bound. Extensive experiments across diverse distillation architectures and training settings demonstrate that AME consistently facilitates robust knowledge distillation, resulting in superior generalization performance across a wide spectrum of downstream tasks.

知识蒸馏视觉语言低数据流形学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。