提出首个视觉多模态并行推理框架,突破传统逐行思考瓶颈。
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
- 采用分块视觉并行推理,通过帕注意力与位置编码提升路径多样性。
- 在V*、CountBench等数据集上显著优于串行推理方法,推理效率提升3倍。
- 适合需要快速高精度视觉理解的多模态应用,如智能客服、医疗影像分析。
现有大模型测试时扩展规律强调通过延长推理长度实现自我反思行为,但垂直扩展常因思维定式导致探索能力受限。本研究转向并行性,提出两种视觉并行推理策略。在此基础上,首次构建适用于多模态大模型的并行推理框架Visual Para-Thinker。为保持路径独立性与推理多样性,引入Pa-Attention与LPRoPE机制。基于vLLM框架实现原生多模态支持,保障高效并行处理。在V*、CountBench、RefCOCO及HallusionBench等基准测试中,结果验证该方法成功将并行推理优势拓展至视觉领域。
原文摘要 · Abstract (English)
Existing LLM test-time scaling laws emphasize the emergence of self-reflective behaviors through extended reasoning length. Nevertheless, this vertical scaling strategy often encounters plateaus in exploration as the model becomes locked into specific thinking pattern. By shifting from depth to parallelism, parallel thinking mitigates the narrowing of exploration. However, the extension of this paradigm to visual domain remains an open research question. In this paper, we first examine the role of visual partitioning in parallelized reasoning and subsequently propose two distinct strategies. Based on the above, we introduce Visual Para-Thinker, representing the inaugural parallel reasoning framework for MLLMs. To maintain path independence and promote diversity in reasoning, our approach integrates Pa-Attention alongside LPRoPE. Leveraging the vLLM framework, we have developed a native multimodal implementation that facilitates high-efficiency parallel processing. Empirical results on benchmark datasets such as V*, CountBench, RefCOCO, and HallusionBench confirm that Visual Para-Thinker successfully extends the benefits of parallel reasoning to the visual domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。