DAMA动态调整多模态模型对难易数据的响应,减少幻觉。
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
- 根据数据难易度和模型实时输出动态优化训练
- 在Object-HalBench上幻觉率降低90%以上
- 适合需要高可信多模态生成的场景
直接偏好优化(DPO)在对齐多模态大语言模型(MLLM)与人类偏好方面表现有效。然而,现有方法对不同难度的数据响应不均,倾向于过拟合易于区分的数据,而忽略难以区分的数据。本文提出数据与模型感知的DPO(DAMA),从两个关键方面动态调整优化过程:(1) 基于数据难易度的数据感知策略;(2) 融合模型实时输出的模型感知策略。通过结合两者,DAMA使模型能够有效适应不同难度的数据。在五个基准上的大量实验表明,DAMA不仅显著提升可信度,还改善了通用任务性能。例如,在Object-HalBench上,DAMA-7B将响应级和提及级幻觉分别降低了90.0%和95.3%,优于GPT-4V。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has shown effectiveness in aligning multi-modal large language models (MLLM) with human preferences. However, existing methods exhibit an imbalanced responsiveness to the data of varying hardness, tending to overfit on the easy-to-distinguish data while underfitting on the hard-to-distinguish data. In this paper, we propose Data- and Model-aware DPO (DAMA) to dynamically adjust the optimization process from two key aspects: (1) a data-aware strategy that incorporates data hardness, and (2) a model-aware strategy that integrates real-time model responses. By combining the two strategies, DAMA enables the model to effectively adapt to data with varying levels of hardness. Extensive experiments on five benchmarks demonstrate that DAMA not only significantly enhances the trustworthiness, but also improves the effectiveness over general tasks. For instance, on the Object-HalBench, our DAMA-7B reduces response-level and mentioned-level hallucination by 90.0% and 95.3%, respectively, surpassing the performance of GPT-4V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。