arXiv:2508.04017cs.CV2025-08被引 4

测试大模型能否主动识别错误输入,发现多数模型需提示才敢挑错。

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability

  • 构建七类错误输入与三指标框架,系统评估模型检错能力。
  • 多数模型无提示时无法主动发现文本错误,逻辑谬误最易识别。
  • 不同模型依赖视觉或文本不一,需提升主动验证输入真实性能力。

大型多模态模型(LMMs)在复杂多模态任务中表现出色,但近期研究指出大语言模型常被动接受有缺陷输入,导致无效推理。然而,LMMs是否能主动检测并审查错误输入仍未知。为此,我们提出输入审查能力评估框架(ISEval),涵盖七类错误前提和三种评估指标。对十种先进LMMs的广泛评估显示:多数模型在无引导时难以主动识别文本错误,高度依赖显式提示;错误类型影响表现——逻辑谬误识别能力强,表面语言错误和特定条件谬误则较弱;模态信任度各异:Gemini 2.5 pro与Claude Sonnet 4能平衡视觉与文本信息,aya-vision-8b则在冲突中过度依赖文本。这些发现凸显亟需增强LMMs主动验证输入有效性的能力,并为缓解该问题提供新视角。代码已开源:https://github.com/MLGroupJLU/LMM_ISEval。

原文摘要 · Abstract (English)

Large Multimodal Models (LMMs) have witnessed remarkable growth, showcasing formidable capabilities in handling intricate multimodal tasks with exceptional performance. Recent research has underscored the inclination of large language models to passively accept defective inputs, often resulting in futile reasoning on invalid prompts. However, the same critical question of whether LMMs can actively detect and scrutinize erroneous inputs still remains unexplored. To address this gap, we introduce the Input Scrutiny Ability Evaluation Framework (ISEval), which encompasses seven categories of flawed premises and three evaluation metrics. Our extensive evaluation of ten advanced LMMs has identified key findings. Most models struggle to actively detect flawed textual premises without guidance, which reflects a strong reliance on explicit prompts for premise error identification. Error type affects performance: models excel at identifying logical fallacies but struggle with surface-level linguistic errors and certain conditional flaws. Modality trust varies-Gemini 2.5 pro and Claude Sonnet 4 balance visual and textual info, while aya-vision-8b over-rely on text in conflicts. These insights underscore the urgent need to enhance LMMs' proactive verification of input validity and shed novel insights into mitigating the problem. The code is available at https://github.com/MLGroupJLU/LMM_ISEval.

多模态模型输入检测模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。