首个系统评估多模态大模型思维链能力的基准,揭示其推理质量与效率瓶颈。
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
- 构建六领域多模态思维链评测集,引入质量、鲁棒性、效率三维度新指标。
- 带反思机制模型(如Kimi k1.5)推理质量最优,但自修正阶段效率极低。
- 思维链提示在视觉任务中反而降低表现,暴露过拟合式思考风险。
思维链(CoT)显著提升了大语言模型的推理能力,但其对大视听模型(LMMs)的影响尚缺乏系统评估。本文提出MME-CoT,首个专注于多模态大模型思维链推理的基准,覆盖数学、科学、OCR、逻辑、时空与通用场景六大领域。通过高质量数据与独特评估策略,构建包含推理质量、鲁棒性与效率的细粒度评估体系。分析显示:1)具备反思机制的模型(如Kimi k1.5)在推理质量上超越GPT-4o;2)CoT提示常损害感知密集型任务表现,暗示过度思考可能有害;3)虽推理质量高,但含反思的模型在响应与自纠错阶段均存在显著效率瓶颈。该工作为推进多模态推理研究提供基础。
原文摘要 · Abstract (English)
Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation. In this paper, we introduce MME-CoT, a specialized benchmark evaluating the CoT reasoning performance of LMMs, spanning six domains: math, science, OCR, logic, space-time, and general scenes. As the first comprehensive study in this area, we propose a thorough evaluation suite incorporating three novel metrics that assess the reasoning quality, robustness, and efficiency at a fine-grained level. Leveraging curated high-quality data and a unique evaluation strategy, we conduct an in-depth analysis of state-of-the-art LMMs, uncovering several key insights: 1) Models with reflection mechanism demonstrate a superior CoT quality, with Kimi k1.5 outperforming GPT-4o and demonstrating the highest quality results; 2) CoT prompting often degrades LMM performance on perception-heavy tasks, suggesting a potentially harmful overthinking behavior; and 3) Although the CoT quality is high, LMMs with reflection exhibit significant inefficiency in both normal response and self-correction phases. We hope MME-CoT serves as a foundation for advancing multimodal reasoning in LMMs. Project Page: https://mmecot.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。