arXiv:2603.18118cs.CVcs.AI2026-03被引 3

让多模态大模型学会复杂长链视觉推理,自动生成训练数据并持续自我优化。

Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models

  • 用多智能体框架自动生成图像与视频的复杂推理路径。
  • 在多个基准上显著提升长时序视觉推理能力,同时保持基础感知性能。
  • 通过反馈循环实现模型持续进化,适合研究长链推理的学者与开发者。

大型语言模型通过测试时的扩展推理取得了卓越性能,但将此能力推广至多模态大语言模型(MLLMs)仍面临高质长链推理数据稀缺与训练流程不完善两大挑战。为此,我们提出统一的多智能体视觉推理框架,从基础的图像中心模型 Insight-V 进化为通用的时空架构 Insight-V++。首先设计可扩展的数据生成流水线,结合多粒度评估,无需人工干预即可在图像与视频领域自动生成结构化、复杂的推理轨迹。鉴于直接监督此类复杂数据效果不佳,我们构建双智能体架构:推理智能体执行深度分析链,摘要智能体批判性评估并提炼最终结果。尽管初始框架采用直接偏好优化(DPO),其离策略特性限制了强化学习潜力,为此,Insight-V++引入两项新算法——ST-GRPO 和 J-GRPO,分别增强时空推理能力与评估鲁棒性。关键在于,利用摘要智能体提供的可靠反馈,驱动迭代推理路径生成,实现整个多智能体系统的持续重训练与自我提升。大量实验表明,在 LLaVA-NeXT 与 Qwen2.5-VL 等基线模型上,该框架在多个具有挑战性的图像与视频推理基准上均取得显著性能提升,同时保持对传统感知任务的强大能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable reliability and advanced capabilities through extended test-time reasoning. However, extending these capabilities to Multi-modal Large Language Models (MLLMs) remains a significant challenge due to a critical scarcity of high-quality, long-chain reasoning data and optimized training pipelines. To bridge this gap, we present a unified multi-agent visual reasoning framework that systematically evolves from our foundational image-centric model, Insight-V, into a generalized spatial-temporal architecture, Insight-V++. We first propose a scalable data generation pipeline equipped with multi-granularity assessment that autonomously synthesizes structured, complex reasoning trajectories across image and video domains without human intervention. Recognizing that directly supervising MLLMs with such intricate data yields sub-optimal results, we design a dual-agent architecture comprising a reasoning agent to execute extensive analytical chains, and a summary agent to critically evaluate and distill final outcomes. While our initial framework utilized Direct Preference Optimization (DPO), its off-policy nature fundamentally constrained reinforcement learning potential. To overcome these limitations, particularly for long-horizon video understanding, Insight-V++ introduces two novel algorithms, ST-GRPO and J-GRPO, which enhance spatial-temporal reasoning and improve evaluative robustness. Crucially, by leveraging reliable feedback from the summary agent, we guide an iterative reasoning path generation process, retraining the entire multi-agent system in a continuous, self-improving loop. Extensive experiments on base models like LLaVA-NeXT and Qwen2.5-VL demonstrate significant performance gains across challenging image and video reasoning benchmarks while preserving strong capabilities on traditional perception-focused tasks.

视觉推理多模态自进化长链思考

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。