arXiv:2501.01346cs.CVcs.CL2025-01EMNLP综述被引 36

剖析大模型图文对齐与错位问题,揭示根源并提出改进方向

Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability

论文配图:Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
图 1 · 摘自论文原文
  • 从可解释性视角分析图文表征对齐机制
  • 发现数据、模型、推理三层面存在错位现象
  • 适合关注大模型可靠性与可解释性的研究者

大型视觉语言模型(LVLMs)在处理视觉与文本信息方面展现出显著能力。然而,视觉与文本表征之间的对齐问题尚未完全理解。本综述通过可解释性视角,全面考察了LVLM中的对齐与错位现象。首先探讨对齐的基本原理,包括表征与行为特征、训练方法及理论基础;随后分析对象、属性和关系三个语义层级上的错位现象。研究发现,错位源于数据、模型和推理三个层面的挑战。本文系统回顾现有缓解策略,将其分为参数冻结与参数调优两类。最后,提出未来研究方向,强调建立标准化评估协议和深入可解释性研究的必要性。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in processing both visual and textual information. However, the critical challenge of alignment between visual and textual representations is not fully understood. This survey presents a comprehensive examination of alignment and misalignment in LVLMs through an explainability lens. We first examine the fundamentals of alignment, exploring its representational and behavioral aspects, training methodologies, and theoretical foundations. We then analyze misalignment phenomena across three semantic levels: object, attribute, and relational misalignment. Our investigation reveals that misalignment emerges from challenges at multiple levels: the data level, the model level, and the inference level. We provide a comprehensive review of existing mitigation strategies, categorizing them into parameter-frozen and parameter-tuning approaches. Finally, we outline promising future research directions, emphasizing the need for standardized evaluation protocols and in-depth explainability studies.

大模型图文对齐可解释性综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。