探索多模态思维链的推理时扩展,提升跨模态推理能力
Investigating Inference-time Scaling for Chain of Multi-modal Thought: A Preliminary Study
- 引入视觉与文本融合的思维链,实现多模态推理
- 多模态思维比纯文本思维表现更优,且思考更多样
- 适合关注多模态模型推理机制的研究者
近期研究表明,推理时扩展思维链(CoT)在多模态推理任务中具有潜力。尽管现有研究多聚焦于文本思维,但将视觉与文本模态融入推理过程仍属空白。本研究首次系统探索推理时扩展的多模态思维链,在10个跨领域的挑战性任务上评估了主流采样和树搜索方法。统一采用增强一致性验证器,以确保不同思维范式下的有效引导。结果表明,多模态思维显著优于传统纯文本思维,且混合模态可激发更丰富的推理路径。然而,处理更丰富的视觉输入导致更高token消耗,对实际应用提出挑战。本研究揭示该方向的优势与局限,旨在启发未来工作。
原文摘要 · Abstract (English)
Recently, inference-time scaling of chain-of-thought (CoT) has been demonstrated as a promising approach for addressing multi-modal reasoning tasks. While existing studies have predominantly centered on text-based thinking, the integration of both visual and textual modalities within the reasoning process remains unexplored. In this study, we pioneer the exploration of inference-time scaling with multi-modal thought, aiming to bridge this gap. To provide a comprehensive analysis, we systematically investigate popular sampling-based and tree search-based inference-time scaling methods on 10 challenging tasks spanning various domains. Besides, we uniformly adopt a consistency-enhanced verifier to ensure effective guidance for both methods across different thought paradigms. Results show that multi-modal thought promotes better performance against conventional text-only thought, and blending the two types of thought fosters more diverse thinking. Despite these advantages, multi-modal thoughts necessitate higher token consumption for processing richer visual inputs, which raises concerns in practical applications. We hope that our findings on the merits and drawbacks of this research line will inspire future works in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。