用偏好优化与集成提升视觉语言模型的推理分割精度
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
- 基于人类偏好优化模型,提升文本与分割结果质量
- 多输出融合采用偏好注意力机制,准确率领先现有方法
- 专为分割任务设计数据收集与损失函数,减少幻觉
现有基于视觉语言模型(LVLM)的推理分割方法常出现分割不精确和文本生成幻觉问题。本文提出POpen框架,通过偏好优化微调LVLM,使其更符合人类偏好,从而生成更优的文本响应与分割结果。同时引入基于偏好的集成推理方法,在推理阶段利用偏好得分注意力机制整合多个模型输出以实现精细化修正。为更好适配分割任务,框架中还包含基于课程学习机制的新颖分割偏好数据收集方法,以及用于增强分割能力的新型偏好优化损失函数。实验表明,该方法在推理分割任务上达到当前最优性能,相比LISA、PixelLM等先进方法,文本幻觉最小,分割准确率最高。
原文摘要 · Abstract (English)
Existing LVLM-based reasoning segmentation methods often suffer from imprecise segmentation results and hallucinations in their text responses. This paper introduces POPEN, a novel framework designed to address these issues and achieve improved results. POPEN includes a preference-based optimization method to finetune the LVLM, aligning it more closely with human preferences and thereby generating better text responses and segmentation results. Additionally, POPEN introduces a preference-based ensemble method for inference, which integrates multiple outputs from the LVLM using a preference-score-based attention mechanism for refinement. To better adapt to the segmentation task, we incorporate several task-specific designs in our POPEN framework, including a new approach for collecting segmentation preference data with a curriculum learning mechanism, and a novel preference optimization loss to refine the segmentation capability of the LVLM. Experiments demonstrate that our method achieves state-of-the-art performance in reasoning segmentation, exhibiting minimal hallucination in text responses and the highest segmentation accuracy compared to previous advanced methods like LISA and PixelLM. Project page is https://lanyunzhu.site/POPEN/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。