arXiv:2601.18231cs.LGcs.AI2026-01中稿 · AISTATS 20226被引 1

提出新框架优化跨模态微调中特征对齐与目标拟合的平衡。

Rethinking Cross-Modal Fine-Tuning: Optimizing the Interaction Between Feature Alignment and Target Fitting

  • 引入特征-标签失真概念,理论分析对齐与拟合的交互机制。
  • 建立可证明的目标误差泛化界,指导实际算法设计。
  • 在多个数据集上显著超越现有方法,适合跨模态学习研究者。

由于跨学科知识融合需求的增长,将预训练模型适配到未见特征模态变得日益重要。核心挑战在于如何将新模态的表示对齐到预训练模型表示空间中最相关部分,以实现准确的知识迁移。这需要结合特征对齐与目标微调,但未经校准的组合可能加剧源与目标特征-标签结构之间的错位,降低目标泛化能力。现有工作缺乏对此关键交互的理论理解。为此,我们构建了一个原则性框架,建立了目标误差的可证明泛化界,通过一个新颖的特征-标签失真概念解释了特征对齐与目标拟合之间的相互作用。该界为实际算法设计提供了可操作的洞见。所提出的方案在广泛基准数据集上显著优于当前最优方法。

原文摘要 · Abstract (English)

Adapting pre-trained models to unseen feature modalities has become increasingly important due to the growing need for cross-disciplinary knowledge integration. A key challenge here is how to align the representation of new modalities with the most relevant parts of the pre-trained model's representation space to enable accurate knowledge transfer. This requires combining feature alignment with target fine-tuning, but uncalibrated combinations can exacerbate misalignment between the source and target feature-label structures and reduce target generalization. Existing work, however, lacks a theoretical understanding of this critical interaction between feature alignment and target fitting. To bridge this gap, we develop a principled framework that establishes a provable generalization bound on the target error, which explains the interaction between feature alignment and target fitting through a novel concept of feature-label distortion. This bound offers actionable insights into how this interaction should be optimized for practical algorithm design. The resulting approach achieves significantly improved performance over state-of-the-art methods across a wide range of benchmark datasets.

跨模态学习微调优化理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。