让视觉模型根据上下文自动选择推理方式,提升通用理解能力。
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- 统一多种推理模式,通过上下文动态选择最佳策略。
- 在多个场景中实现一致性能提升,显著增强模型泛化能力。
- 适合需要灵活应对复杂视觉任务的研究与应用。
当前视觉推理方法多聚焦于特定推理模式,虽在特定领域有改进,但难以发展通用推理能力。受此启发,我们提出一种新的自适应推理范式——视觉思维混合模型(Mixture-of-Visual-Thoughts, MoVT),将多种推理模式统一于单一模型中,并根据上下文引导其选择合适模式。为此,我们设计了两阶段的自适应视觉推理学习框架AdaVaR:在监督冷启动阶段统一学习不同模式,再通过强化学习过程结合精心设计的AdaGRPO算法,诱导模型具备模式选择能力。大量实验表明,AdaVaR能有效引导模型学习并区分多种模式,实现上下文自适应的模式选择,在多种场景下均取得稳定提升,验证了MoVT作为构建通用视觉推理模型的有效方案。
原文摘要 · Abstract (English)
Current visual reasoning methods mainly focus on exploring specific reasoning modes. Although improvements can be achieved in particular domains, they struggle to develop general reasoning capabilities. Inspired by this, we propose a novel adaptive reasoning paradigm, Mixture-of-Visual-Thoughts (MoVT), which unifies different reasoning modes within a single model and guides it to select the appropriate mode based on context. To achieve this, we introduce AdaVaR, a two-stage Adaptive Visual Reasoning learning framework: different modes are unified and learned during the supervised cold-start stage, and the mode selection capability is induced via an RL process with a carefully designed AdaGRPO algorithm. Extensive experiments show that AdaVaR effectively guides the model to learn and differentiate multiple modes and perform context-adaptive mode selection, achieving consistent improvement across various scenarios, highlighting MoVT as an effective solution for building general visual reasoning models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。