arXiv:2412.17417cs.CV2024-12被引 5

用合成数据提升视觉语言模型可信度,显著降低幻觉率。

Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment

  • 用奖励模型模拟人类偏好生成合成数据用于对齐训练
  • 使模型幻觉率下降50.98%,在多个评测中表现大幅提升
  • 适合关注模型安全性与真实应用的开发者和研究者

大型视觉语言模型(LVLMs)通过融合视觉与文本信息展现出强大的理解与生成能力。然而,当前模型仍易产生幻觉,严重影响实际应用中的性能与用户体验。后训练对齐,特别是偏好微调(preference-tuning),旨在使模型输出与行为(安全、指令遵循、风格)一致,提升任务适应性与鲁棒性。现有方法多依赖强模型或基础模型(如CLIP)判断图像-文本对的优劣,但多模态场景下的合成数据对齐仍缺乏探索。本文提出SynthAlign,一个专为基于直接偏好优化(DPO)的后训练对齐设计的合成人类偏好数据生成与收集流程。核心在于使用奖励模型作为人类偏好的代理。通过一系列评估与基准测试验证框架有效性及数据集价值。结果显示,经该框架优化的LLaVA-1.5-7B模型在POPE任务中准确率达到87.6%,精确率达97.8%;MMHal-Bench得分从2.36提升至3.49;幻觉率由51.0%降至25.0%(相对减少50.98%)。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have shown promising capabilities in understanding and generating information by integrating both visual and textual data. However, current models are still prone to hallucinations, which degrade the performance and greatly harm the user experience in real-world applications. Post-training alignment, particularly preference-tuning, is intended to align model outputs and behaviors (safety, instruction-following, style), ensuring robustness and adaptability to a wide range of tasks. The use of synthetic data for alignment, particularly in multimodal settings, remains under explored. Existing approaches typically use a strong model or a ground-truth model (CLIP) to determine positive and negative image-text data points. This paper proposes SynthAlign, a pipeline to generate and collect synthetic human-preference image-text data with optimal control built specifically for post-training alignment with DPO. At the core of the framework is the utilization of reward models as a proxy of human preference. A series of evaluation and benchmarking is provided to validate the effectiveness of the proposed framework and the resulting dataset. Notably, our framework enhanced LLaVA-1.5-7B achieved substantial POPE improvements: 87.6\% accuracy and 97.8\% precision, MMHal-Bench score increased from 2.36 to 3.49, and hallucination rate decreased from 51.0\% to 25.0\% (a 50.98\% relative reduction).

视觉语言模型偏好对齐幻觉抑制合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。