arXiv:2603.00720cs.LG2026-03

MARS自动寻找多模态模型最优低秩适配参数,提升训练效率与准确率。

MARS: Harmonizing Multimodal Convergence via Adaptive Rank Search

  • 用双尺度规律动态调节不同模态的训练速度,实现自适应优化。
  • 在多个数据集上显著提升性能,最高达12.3%的准确率增益。
  • 适合需要高效微调多模态大模型的研究者和工程师使用。

使用参数高效方法(如低秩适配,LoRA)微调多模态大语言模型(MLLM)对任务适配至关重要。然而,各模态间不平衡的训练动态常导致负向干扰,降低准确率,传统方法依赖手动调节学习率等低效启发式策略。为此,我们提出MARS(Multimodal Adaptive Rank Search),通过发现最优秩组合来平衡训练动态并最大化性能。核心创新在于双尺度规律框架:其一建模模块特异性收敛时间,以修剪搜索空间至动态对齐候选;其二预测最终任务性能,从修剪后集合中选出最优秩对。通过将LoRA秩重用于调控模态特异性收敛速度,MARS优于基线方法,提供一种鲁棒且自动化的MLLM微调优化策略。

原文摘要 · Abstract (English)

Fine-tuning Multimodal Large Language Models (MLLMs) with parameter-efficient methods like Low-Rank Adaptation (LoRA) is crucial for task adaptation. However, imbalanced training dynamics across modalities often lead to suboptimal accuracy due to negative interference, a challenge typically addressed with inefficient heuristic methods such as manually tuning separate learning rates. To overcome this, we introduce MARS (Multimodal Adaptive Rank Search), an approach to discover optimal rank pairs that balance training dynamics while maximizing performance. Our key innovation, a proposed framework of dual scaling laws, enables this search: one law models module-specific convergence time to prune the search space to candidates with aligned dynamics, while the other predicts final task performance to select the optimal pair from the pruned set. By re-purposing the LoRA rank as a controller for modality-specific convergence speed, MARS outperforms baseline methods and provides a robust, automated strategy for optimizing MLLM fine-tuning.

多模态低秩适配自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。