首个系统评估多模态大模型诚实性的基准,测试其面对无法回答的视觉问题时的表现。
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
- 构建包含1.2万+样本的诚实性评测基准MoHoBench,涵盖四类不可答视觉问题。
- 28个主流多模态大模型中多数无法在应拒绝时正确拒答,诚实性普遍不足。
- 发现模型诚实性受视觉信息影响显著,需专门方法对齐多模态诚实行为。
近期多模态大语言模型(MLLMs)在视觉-语言任务上取得显著进展,但可能生成有害或不可信内容。尽管已有大量研究关注语言模型的可信性,但针对多模态大模型在面对视觉不可答问题时是否诚实的行为仍缺乏系统探讨。本文首次系统评估了多种MLLMs在诚实性方面的表现,将诚实性定义为模型对不可答视觉问题的回应行为,提出四类代表性问题类型,并构建了大规模多模态诚实性基准MoHoBench,包含超过1.2万张图像-问题样本,通过多阶段筛选与人工验证确保质量。基于MoHoBench,我们对28个主流多模态大模型进行了评测与分析。结果表明:(1) 多数模型在应拒绝回答时未能正确拒绝;(2) 模型的诚实性不仅涉及语言建模,更深度依赖视觉信息,亟需专用方法实现多模态诚实对齐。为此,我们初步采用监督学习和偏好学习实现了诚实性对齐,为未来可信多模态大模型研究奠定基础。数据与代码详见https://github.com/yanxuzhu/MoHoBench。
原文摘要 · Abstract (English)
Recently Multimodal Large Language Models (MLLMs) have achieved considerable advancements in vision-language tasks, yet produce potentially harmful or untrustworthy content. Despite substantial work investigating the trustworthiness of language models, MMLMs' capability to act honestly, especially when faced with visually unanswerable questions, remains largely underexplored. This work presents the first systematic assessment of honesty behaviors across various MLLMs. We ground honesty in models' response behaviors to unanswerable visual questions, define four representative types of such questions, and construct MoHoBench, a large-scale MMLM honest benchmark, consisting of 12k+ visual question samples, whose quality is guaranteed by multi-stage filtering and human verification. Using MoHoBench, we benchmarked the honesty of 28 popular MMLMs and conducted a comprehensive analysis. Our findings show that: (1) most models fail to appropriately refuse to answer when necessary, and (2) MMLMs' honesty is not solely a language modeling issue, but is deeply influenced by visual information, necessitating the development of dedicated methods for multimodal honesty alignment. Therefore, we implemented initial alignment methods using supervised and preference learning to improve honesty behavior, providing a foundation for future work on trustworthy MLLMs. Our data and code can be found at https://github.com/yanxuzhu/MoHoBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。