不用模型当裁判,让多模态大模型高效自我改进。
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
- 用可控幻觉生成对比数据对,避免依赖模型自评。
- 在多个基准上实现更高精度与召回率,计算成本更低。
- 适合追求高效迭代的多模态模型研发团队。
多模态大语言模型(MLLMs)的自提升对于提高其可靠性与鲁棒性至关重要。然而,现有方法通常严重依赖MLLM自身作为评判者,导致计算开销高,并可能引发奖励欺骗和模型坍塌等问题。本文提出一种新型的、无需模型评判的自提升框架。该方法通过可控幻觉机制生成偏好学习样本对,并利用轻量级对比式图文编码器评估与反转数据对,以优化数据质量。在公开基准及新提出的挑战幻觉控制的IC数据集上的评估表明,该模型在保持更低成本的同时,优于传统技术,在精度与召回率方面均有提升。该方法为可扩展的MLLM自提升提供了高效路径,兼顾性能提升与资源节约。
原文摘要 · Abstract (English)
Self-improvement in multimodal large language models (MLLMs) is crucial for enhancing their reliability and robustness. However, current methods often rely heavily on MLLMs themselves as judges, leading to high computational costs and potential pitfalls like reward hacking and model collapse. This paper introduces a novel, model-level judge-free self-improvement framework. Our approach employs a controlled feedback mechanism while eliminating the need for MLLMs in the verification loop. We generate preference learning pairs using a controllable hallucination mechanism and optimize data quality by leveraging lightweight, contrastive language-image encoders to evaluate and reverse pairs when necessary. Evaluations across public benchmarks and our newly introduced IC dataset designed to challenge hallucination control demonstrate that our model outperforms conventional techniques. We achieve superior precision and recall with significantly lower computational demands. This method offers an efficient pathway to scalable self-improvement in MLLMs, balancing performance gains with reduced resource requirements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。