arXiv:2508.14264cs.CV2025-08NeurIPS被引 5

通过排序重建提升视觉文本对齐,增强大模型鲁棒性。

Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models

  • 引入图像与文本顺序重建任务,强化跨模态对齐。
  • 在多个基准上达到当前最优性能,显著提升理解能力。
  • 适合需要强视觉推理的多模态应用开发者参考。

大型多模态模型(LMMs)因在各类理解任务中表现优异而备受关注,但仍存在鲁棒性和泛化能力不足的问题,根源在于视觉与文本特征间的对齐与相关性。本文提出一种简单高效的训练机制,通过解决特征打乱问题来改善视觉-文本模态间的鲁棒对齐。具体方法是在预训练和微调阶段引入两项新任务:重构图像顺序与文本顺序。同时,提出定向标记(Directed-Tokens)方法以捕捉视觉与文本知识,支持正确重建视觉输入顺序;并设计图像到响应引导损失(Image-to-Response Guided Loss),进一步提升模型在生成回应时的视觉理解能力。所提方法在学术任务导向及指令遵循型LMM基准上持续取得最先进性能。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) have gained impressive performance due to their outstanding capability in various understanding tasks. However, these models still suffer from some fundamental limitations related to robustness and generalization due to the alignment and correlation between visual and textual features. In this paper, we introduce a simple but efficient learning mechanism for improving the robust alignment between visual and textual modalities by solving shuffling problems. In particular, the proposed approach can improve reasoning capability, visual understanding, and cross-modality alignment by introducing two new tasks: reconstructing the image order and the text order into the LMM's pre-training and fine-tuning phases. In addition, we propose a new directed-token approach to capture visual and textual knowledge, enabling the capability to reconstruct the correct order of visual inputs. Then, we introduce a new Image-to-Response Guided loss to further improve the visual understanding of the LMM in its responses. The proposed approach consistently achieves state-of-the-art (SoTA) performance compared with prior LMMs on academic task-oriented and instruction-following LMM benchmarks.

多模态对齐视觉理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。