发现多模态大模型常误选错误答案,无法识别正确选项不存在。
When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding

- 测试模型在无正确选项时能否识别出无解
- 多数模型仍选干扰项,尤其在时间推理任务中更差
- 思维链提示能改善但仍不足够,需新机制
多模态大语言模型(MLLMs)在视频理解上取得显著进展,但其回答可靠性仍缺乏深入研究。本文针对视频理解中‘正确答案缺失’的检测问题开展诊断分析:当正确答案不在候选选项中时,可靠模型应能识别出‘无有效选项’。我们在三种场景下评估:含‘以上皆非’选项的多选题、带检测指令的开放式生成、以及无任何引导的标准评测。在多种模型与基准数据集上,结果显示大多数模型仍倾向于选择看似合理的干扰项,而非识别答案缺失。这一缺陷在时间推理任务中尤为严重,且随着帧采样密度增加而加剧。我们进一步探索思维链提示作为缓解策略,发现虽能显著提升检测率,但整体表现仍不理想,表明仅靠提示方法不足以解决根本问题。研究揭示了多模态系统在缺失答案检测上的系统性失败,强调必须引入显式检测机制。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have made substantial advancements in video understanding, yet the reliability of their responses remains underexplored. This work presents a diagnostic study of absent answer detection for MLLMs in video understanding, where the correct answer is deliberately excluded from the candidate set and a reliable model is expected to recognize that no valid option exists. We evaluate the absent answer detection behavior under three settings: multiple-choice questions augmented with an ``None of the Above'' option, open-ended generation with a detection instruction, and standard evaluation without any guidance. Across a diverse set of models and benchmarks, we find that MLLMs overwhelmingly select plausible distractors rather than detecting the absent answer. This failure is more pronounced in temporal reasoning tasks and worsens with denser frame sampling. We further explore chain-of-thought prompting as a mitigation strategy and find that while it substantially improves detection rates, performance remains unsatisfactory, suggesting that prompting-based strategies alone are insufficient to fully address this limitation. These findings expose a systematic failure in absent answer detection and highlight the need for explicit detection mechanisms in multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。