揭示多模态搜索中的隐性失败,提出诊断框架并验证其普遍性
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

- 构建六类故障分类体系,涵盖模态捷径、幻觉溯源等隐蔽问题
- 实测显示表面准确率严重高估真实轨迹正确率,最高偏差达40%
- 适用于评估多模态大模型的可靠性,尤其关注工具调用链的可信度
多模态智能体搜索系统依赖外部工具回答知识密集型视觉问题。现有评估主要关注最终答案准确性,可能忽略搜索路径中的隐藏失效。本文研究此类隐性可靠性问题——沉默失败。提出一个涵盖模态捷径、幻觉定位、错误证据正确答案、过度检索清洗、跨模态矛盾和溯源幻觉的六类分类体系。基于此,构建轨迹级诊断流水线,在统一ReAct框架下评估答案正确性与证据可追溯性。在MMSearch-Plus数据集上对四个前沿多模态模型的轨迹进行实验,结果表明表面准确率持续高估真实轨迹正确率。通过跨评审验证、空白图像压力测试和工具消融分析进一步证明,沉默失败具有能力依赖性,常发生转移而非消失。
原文摘要 · Abstract (English)
Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing evaluations mainly focus on final-answer accuracy and may miss failures in the search trajectory. In this work, we study such hidden reliability issues as silent failures. We introduce a six-category taxonomy covering modality shortcuts, phantom grounding, wrong-evidence-right-answer cases, over-retrieval laundering, cross-modal contradiction, and provenance hallucination. Based on this taxonomy, we build a trajectory-level diagnostic pipeline that evaluates both answer correctness and evidence-grounding quality under a unified ReAct-style scaffold. Experiments on MMSearch-Plus trajectories across four frontier multimodal models show that surface accuracy consistently overestimates true trajectory-level correctness. We further use cross-judge validation, blank-image stress tests, and tool ablations to show that silent failures are capability-dependent and often shift rather than disappear. Home-page: https://github.com/DingWu1021/silent-failures-multimodal-agentic-search
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。