arXiv:2503.18320cs.AIcs.CL2025-03

让视觉指令更像大模型原生写法,提升多模态模型表现

Bridging Writing Manner Gap in Visual Instruction Tuning by Creating LLM-aligned Instructions

  • 用基座大模型生成与自身文风一致的视觉指令
  • 在15个评测上显著减少幻觉并提升综合性能
  • 适合关注多模态对齐与生成质量的研究者

在大型多模态模型(LMMs)中,视觉指令调优阶段的指令质量显著影响模态对齐效果。本文从独特视角——写作风格(Writing Manner)出发,考察词汇选择、语法与句式结构对语义表达的影响。我们发现,视觉指令与基座大语言模型(LLMs)之间存在显著的写作风格差距,导致预训练基座模型偏离原有写作风格,进而引发其及整个LMM能力下降。为弥合这一差距并保持原始语义,我们提出直接利用基座大模型,将软格式视觉指令的写作风格对齐至其自身风格,生成新型的LLM对齐指令。人工评估表明,该方法有效缩小了写作风格差距。使用这些对齐指令后,基线模型LLaVA-7B和QwenVL在全部15个视觉与语言基准上均表现出更强的抗幻觉能力,并实现非平凡的全面改进。

原文摘要 · Abstract (English)

In the realm of Large Multi-modal Models (LMMs), the instruction quality during the visual instruction tuning stage significantly influences the performance of modality alignment. In this paper, we assess the instruction quality from a unique perspective termed \textbf{Writing Manner}, which encompasses the selection of vocabulary, grammar and sentence structure to convey specific semantics. We argue that there exists a substantial writing manner gap between the visual instructions and the base Large Language Models (LLMs) within LMMs. This gap forces the pre-trained base LLMs to deviate from their original writing styles, leading to capability degradation of both base LLMs and LMMs. To bridge the writing manner gap while preserving the original semantics, we propose directly leveraging the base LLM to align the writing manner of soft-format visual instructions with that of the base LLM itself, resulting in novel LLM-aligned instructions. The manual writing manner evaluation results demonstrate that our approach successfully minimizes the writing manner gap. By utilizing LLM-aligned instructions, the baseline models LLaVA-7B and QwenVL demonstrate enhanced resistance to hallucinations and non-trivial comprehensive improvements across all $15$ visual and language benchmarks.

多模态对齐指令调优大模型风格幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。