新基准CoMT让视觉语言模型像人一样边看边思考,生成视觉操作。
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
- 要求模型输出包含视觉操作的多模态推理过程
- 涵盖创建、删除、更新、选择四类复杂视觉操作
- 适合研究多模态推理与生成的学者参考
大型视觉语言模型(LVLMs)在多模态任务中取得显著进展,尤其在多模态思维链(MCoT)推理方面。然而,现有基准仍采用传统范式:多模态输入、文本输出,导致缺失视觉操作和表达模糊等问题。为此,我们提出新型多模态思维链(CoMT)基准,突破传统限制。CoMT要求输入为多模态,输出也需包含多模态推理,旨在模拟人类整合视觉操作的自然推理过程。该基准包含四大类别:(1) 视觉创建,(2) 视觉删除,(3) 视觉更新,(4) 视觉选择,全面覆盖真实场景中的复杂视觉操作与简洁表达。我们在CoMT上评估多种LVLM及策略,揭示当前方法的能力与局限。期望CoMT能推动更多研究将多模态生成引入推理流程。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks still follow a traditional paradigm with multi-modal input and text-modal output, which leads to significant drawbacks such as missing visual operations and vague expressions. Motivated by this, we introduce a novel Chain of Multi-modal Thought (CoMT) benchmark to address these limitations. Different from the traditional MCoT benchmark, CoMT requires both multi-modal input and multi-modal reasoning output, aiming to mimic human-like reasoning that inherently integrates visual operation. Specifically, CoMT consists of four categories: (1) Visual Creation, (2) Visual Deletion, (3) Visual Update, and (4) Visual Selection to comprehensively explore complex visual operations and concise expression in real scenarios. We evaluate various LVLMs and strategies on CoMT, revealing some key insights into the capabilities and limitations of the current approaches. We hope that CoMT can inspire more research on introducing multi-modal generation into the reasoning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。