arXiv:2510.23479cs.CV2025-10被引 13

用混合图像增强提升多模态模型对齐效果,兼顾效率与泛化。

MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding

  • 基于注意力图聚类生成上下文对齐的混合图像。
  • 在分类任务中准确率超越传统方法,多模态对齐能力显著提升。
  • 适合追求高效稳定训练的多模态模型开发者使用。

多模态大语言模型(MLLMs)在后训练阶段的视觉-语言对齐依赖监督微调(SFT)或强化学习(RL)。SFT稳定但需人工标注且缺乏任务泛化性,而RL虽能通过奖励信号优化答案,却存在计算开销大和不稳定的缺陷。为平衡可扩展性、效率与对齐泛化性,我们提出MergeMix,一种统一的增强范式,通过基于标记合并的Mixup增强实现SFT与RL的融合。具体地,根据合并后的注意力图聚类区域生成语义对齐的混合图像及对应标签;进一步构建原始图像与MergeMix生成图像之间的偏好对,采用混合SimPO损失优化软偏好边界。大量实验表明,MergeMix不仅在分类任务中达到领先准确率,还显著提升MLLMs的泛化能力和对齐效果,为偏好对齐提供兼具训练效率与稳定性的新范式。

原文摘要 · Abstract (English)

Vision-language alignment in multi-modal large language models (MLLMs) relies on supervised fine-tuning (SFT) or reinforcement learning (RL). To align multi-modal large language models (MLLMs) in the post-training stage, supervised fine-tuning (SFT) is a stable choice but requires human annotations and lacks task generalizations, while Reinforcement Learning (RL) searches for better answers from reward signals but suffers from computational overhead and instability. To achieve balance among scalability, efficiency, and alignment generalizations, we propose MergeMix, a unified paradigm that bridges SFT and RL with an efficient Token Merge based Mixup augmentation. As for the Mixup policy, we generate contextual aligned mixed images with the corresponding labels according to the merged attention maps with cluster regions. Then, we enhance the preference-driven paradigm for MLLMs by building preference pairs with raw images and MergeMix-generated ones and optimizing the soft preference margin with the mixed SimPO loss. Extensive experiments demonstrate that MergeMix not only achieves dominant classification accuracy as an augmentation method but also improves generalization abilities and alignment of MLLMs, providing a new learning paradigm for preference alignment with training efficiency and stability.

多模态图像增强对齐学习训练效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。