arXiv:2503.14504cs.CV2025-03综述被引 19

系统梳理多模态大模型对齐人类偏好的方法与进展

Aligning Multimodal LLM with Human Preference: A Survey

  • 从应用、数据、评测三方面构建对齐算法体系
  • 涵盖图像、视频、音频等多场景的对齐方案
  • 适合关注多模态模型安全与可控性的研究者

大语言模型(LLMs)可通过简单提示完成多种通用任务,无需特定训练。基于LLMs的多模态大语言模型(MLLMs)在处理视觉、听觉和文本等复杂任务上展现出巨大潜力。然而,真实性、安全性、类O1推理能力以及与人类偏好的对齐问题仍亟待解决。这催生了多种对齐算法,针对不同应用场景和优化目标。本文旨在全面系统地回顾MLLM对齐算法,重点探讨四个方面:(1) 对齐算法的应用场景,包括通用图像理解、多图、视频、音频及扩展多模态应用;(2) 构建对齐数据集的核心因素,如数据来源、模型输出和偏好标注;(3) 评估对齐算法所用基准;(4) 对未来发展方向的讨论。本工作旨在帮助研究者梳理领域进展并激发更优对齐方法的发展。论文项目页见:https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Alignment。

原文摘要 · Abstract (English)

Large language models (LLMs) can handle a wide variety of general tasks with simple prompts, without the need for task-specific training. Multimodal Large Language Models (MLLMs), built upon LLMs, have demonstrated impressive potential in tackling complex tasks involving visual, auditory, and textual data. However, critical issues related to truthfulness, safety, o1-like reasoning, and alignment with human preference remain insufficiently addressed. This gap has spurred the emergence of various alignment algorithms, each targeting different application scenarios and optimization goals. Recent studies have shown that alignment algorithms are a powerful approach to resolving the aforementioned challenges. In this paper, we aim to provide a comprehensive and systematic review of alignment algorithms for MLLMs. Specifically, we explore four key aspects: (1) the application scenarios covered by alignment algorithms, including general image understanding, multi-image, video, and audio, and extended multimodal applications; (2) the core factors in constructing alignment datasets, including data sources, model responses, and preference annotations; (3) the benchmarks used to evaluate alignment algorithms; and (4) a discussion of potential future directions for the development of alignment algorithms. This work seeks to help researchers organize current advancements in the field and inspire better alignment methods. The project page of this paper is available at https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Alignment.

多模态对齐大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。