arXiv:2412.11196cs.CLcs.CV2024-12被引 1

让多模态大模型学会说‘不知道’,提升可信度

Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal

  • 基于信息边界定义何时该拒绝回答
  • 拒绝准确率提升,且不影响模型帮助性
  • 适合关注模型可靠性与安全性的研究者

多模态大语言模型(MLLMs)在多模态感知与理解方面表现优异,但其生成幻觉或不准确回应的倾向削弱了可信度。现有方法大多忽视了拒绝回答作为提升可靠性的重要手段。为此,本文提出信息边界感知学习框架(InBoL),使MLLMs在信息不足时能主动拒绝回答。据我们所知,InBoL是首个系统性定义MLLM拒绝条件的框架,引入数据生成管道与定制训练策略,显著增强模型做出恰当拒绝的能力。为评估可信度,我们进一步提出以用户为中心的对齐目标及相应指标。实验表明,拒绝准确率显著提升,且未明显影响模型帮助性,确立了InBoL在构建更可信MLLMs中的关键作用。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) excel at multimodal perception and understanding, yet their tendency to generate hallucinated or inaccurate responses undermines their trustworthiness. Existing methods have largely overlooked the importance of refusal responses as a means of enhancing MLLMs reliability. To bridge this gap, we present the Information Boundary-aware Learning Framework (InBoL), a novel approach that empowers MLLMs to refuse to answer user queries when encountering insufficient information. To the best of our knowledge, InBoL is the first framework that systematically defines the conditions under which refusal is appropriate for MLLMs using the concept of information boundaries proposed in our paper. This framework introduces a comprehensive data generation pipeline and tailored training strategies to improve the model's ability to deliver appropriate refusal responses. To evaluate the trustworthiness of MLLMs, we further propose a user-centric alignment goal along with corresponding metrics. Experimental results demonstrate a significant improvement in refusal accuracy without noticeably compromising the model's helpfulness, establishing InBoL as a pivotal advancement in building more trustworthy MLLMs.

多模态可信度拒绝机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。