arXiv:2510.10965cs.CLcs.AI2025-10被引 4

构建新数据集与框架,提升多模态模型识别问题假前提的能力

Judge Before Answer: Can MLLM Discern the False Premise in Question?

  • 按认知能力将假前提分为三类十三种,自动化构建新数据集
  • 现有模型在假前提识别上表现不佳,准确率仍很低
  • 提出针对性增强框架,显著提升模型识别假前提的鲁棒性

近年来,多模态大语言模型(MLLMs)取得了惊人进展。然而,这些模型仍易受假前提问题困扰。现有评测基准覆盖面不足、细粒度分类缺失,难以严格评估模型识别假前提的能力。为此,我们提出一种全自动管道,系统构建全面的假前提问题评测基准。方法依据识别所需能力,将前提细分为三类十三种子类,形成JBA数据集。实验表明,当前MLLMs在假前提识别任务中仍表现不佳。基于该基准,我们进一步提出一种识别增强框架,专门提升模型对假前提的鲁棒性。大量实验验证,使用该框架训练的模型在假前提识别上取得显著提升。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have witnessed astonishing advancements in recent years. Despite these successes, MLLMs remain vulnerable to flase premise problems. However, existing benchmarks targeting this issue are limited in scope: they often lack fine-grained categorization, exhibit insufficient coverage, and thus fail to provide a rigorous evaluation of the ability of models to recognize false premises. To bridge this gap, we introduce a fully automated pipeline for constructing a comprehensive benchmark of false premise questions. Our method systematically categorizes the premises into three main types and thirteen subtypes according to the abilities required to identify the premises, resulting in the JBA dataset.Results show current MLLMs still struggle with false premise recognition. Building upon this benchmark, we further propose a recognition enhancement framework tailored to strengthen the robustness of MLLMs to detect false premises. Extensive experiments demonstrate that models trained with our framework achieve significant improvements in false premise recognition.

多模态模型假前提识别评测基准模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。