arXiv:2411.14432cs.CV2024-11CVPR被引 146

构建长链视觉推理数据与多智能体训练框架,提升多模态模型推理能力。

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

  • 设计两阶段生成与多粒度评估,自动生成高质量长链推理数据
  • 在多个视觉推理基准上显著提升性能,优于基线模型
  • 多智能体系统适配感知任务,兼具推理与泛化优势

大型语言模型通过增强推理能力展现更优表现,从思维链提示发展到产品级解决方案如OpenAI o1。尽管已有诸多改进尝试,视觉语言任务中高质量长链推理数据及优化训练流程仍缺乏深入探索。本文提出Insight-V,首次系统性地实现:1)可扩展生成复杂多模态任务的长而稳健的推理数据;2)有效训练流程以增强多模态大模型(MLLM)的推理能力。为无人工干预生成多样且结构化的长推理路径,设计两步式渐进生成策略,并采用多粒度评估方法保证数据质量。观察发现,直接用此类复杂数据监督MLLM无法获得理想推理效果。为此,构建包含推理代理与摘要代理的多智能体系统,其中推理代理专注长链推理,摘要代理负责评判与总结结果,并引入迭代DPO算法提升生成稳定性与质量。基于流行的LLaVA-NeXT模型及更强基线模型,实验证明Insight-V在多个需视觉推理的挑战性基准上取得显著性能提升。得益于多智能体架构,Insight-V在感知类多模态任务上亦能保持或提升性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate enhanced capabilities and reliability by reasoning more, evolving from Chain-of-Thought prompting to product-level solutions like OpenAI o1. Despite various efforts to improve LLM reasoning, high-quality long-chain reasoning data and optimized training pipelines still remain inadequately explored in vision-language tasks. In this paper, we present Insight-V, an early effort to 1) scalably produce long and robust reasoning data for complex multi-modal tasks, and 2) an effective training pipeline to enhance the reasoning capabilities of multi-modal large language models (MLLMs). Specifically, to create long and structured reasoning data without human labor, we design a two-step pipeline with a progressive strategy to generate sufficiently long and diverse reasoning paths and a multi-granularity assessment method to ensure data quality. We observe that directly supervising MLLMs with such long and complex reasoning data will not yield ideal reasoning ability. To tackle this problem, we design a multi-agent system consisting of a reasoning agent dedicated to performing long-chain reasoning and a summary agent trained to judge and summarize reasoning results. We further incorporate an iterative DPO algorithm to enhance the reasoning agent's generation stability and quality. Based on the popular LLaVA-NeXT model and our stronger base MLLM, we demonstrate significant performance gains across challenging multi-modal benchmarks requiring visual reasoning. Benefiting from our multi-agent system, Insight-V can also easily maintain or improve performance on perception-focused multi-modal tasks.

多模态推理长链思维多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。