用渐进式知识蒸馏提升多模态模型幻觉检测能力
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
- 分层设计蒸馏流程,从粗到细逐步对齐视觉语言知识
- 在多个数据集上显著优于基线模型,提升幻觉检测准确率
- 适合关注多模态模型可信性与鲁棒性的研究者
随着视觉语言模型(VLMs)的快速发展,大型多模态模型的负责任行为成为重要研究方向,尤其聚焦于幻觉检测与事实性验证。本文针对负责任AI挑战的两个赛道提出解决方案。受通用领域启发,更小的蒸馏后VLM常能超越直接微调的大模型,且效率更高。因此,我们从知识蒸馏视角联合解决两项任务,提出一种渐进式混合知识蒸馏框架HKD4VLM。该框架包含金字塔式渐进在线蒸馏与三元耦合精炼蒸馏,层次化地实现从粗粒度知识对齐到细粒度优化。此外,引入映射偏移增强推理与多样化增强策略以提升性能与鲁棒性。大量实验表明HKD4VLM有效,消融研究揭示了关键设计对性能提升的作用。
原文摘要 · Abstract (English)
Driven by the rapid progress in vision-language models (VLMs), the responsible behavior of large-scale multimodal models has become a prominent research area, particularly focusing on hallucination detection and factuality checking. In this paper, we present the solution for the two tracks of Responsible AI challenge. Inspirations from the general domain demonstrate that a smaller distilled VLM can often outperform a larger VLM that is directly tuned on downstream tasks, while achieving higher efficiency. We thus jointly tackle two tasks from the perspective of knowledge distillation and propose a progressive hybrid knowledge distillation framework termed HKD4VLM. Specifically, the overall framework can be decomposed into Pyramid-like Progressive Online Distillation and Ternary-Coupled Refinement Distillation, hierarchically moving from coarse-grained knowledge alignment to fine-grained refinement. Besides, we further introduce the mapping shift-enhanced inference and diverse augmentation strategies to enhance model performance and robustness. Extensive experimental results demonstrate the effectiveness of our HKD4VLM. Ablation studies provide insights into the critical design choices driving performance gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。