arXiv:2604.14041cs.CV2026-04ACL被引 1

评测大模型在日常场景中找关键视觉线索的推理能力

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios

论文配图:Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios
图 1 · 摘自论文原文
  • 构建真实生活场景下的视觉线索推理任务,要求模型主动寻找关键信息
  • 覆盖4大生活领域16个子任务,问题设计超越表面识别
  • 揭示准确识别视觉线索是可靠推理的核心,适合研究多模态推理的学者

日常生活场景具有视觉丰富性,要求多模态大语言模型(MLLMs)在噪声中筛选出决定性视觉线索以实现准确推理。然而,现有基准大多评估模型的先验知识或感知理解能力,忽视了推理这一关键能力。为此,我们提出DailyClue,一个面向日常场景中视觉线索驱动推理的基准。其构建遵循两大原则:(1) 严格基于真实日常活动;(2) 设计具有挑战性的提问,要求超越表层感知。问题不依赖简单识别,而是迫使模型主动探索并利用合适的视觉线索进行后续推理。我们构建了一个涵盖四大主要日常领域和16个不同子任务的综合性数据集。对MLLMs及代理模型的全面评估表明,该基准带来了显著挑战。分析揭示关键洞察:准确识别视觉线索是实现鲁棒推理的关键。

原文摘要 · Abstract (English)

Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual clues for accurate reasoning. Yet, current benchmarks predominantly aim at evaluating MLLMs' pre-existing knowledge or perceptual understanding, often neglecting the critical capability of reasoning. To bridge this gap, we introduce DailyClue, a benchmark designed for visual clue-driven reasoning in daily scenarios. Our construction is guided by two core principles: (1) strict grounding in authentic daily activities, and (2) challenging query design that necessitates more than surface-level perception. Instead of simple recognition, our questions compel MLLMs to actively explore suitable visual clues and leverage them for subsequent reasoning. To this end, we curate a comprehensive dataset spanning four major daily domains and 16 distinct subtasks. Comprehensive evaluation across MLLMs and agentic models underscores the formidable challenge posed by our benchmark. Our analysis reveals several critical insights, emphasizing that the accurate identification of visual clues is essential for robust reasoning.

多模态推理视觉线索日常场景基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。