让视觉思考更高效精准,动态插入图像降低72.6%计算开销
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
- 动态插入视觉信息,按需调用减少冗余
- 视觉推理更连贯,实现72.6%的令牌消耗下降
- 适合追求高效多模态推理的应用场景
近期,跨模态思维链(ICoT)通过结合多模态输入与输出取得了显著进展。然而现有方法仍存在两大局限:(1)视觉思考位置固定,导致推理效率低且不灵活;(2)视觉表示断裂,语义不连贯。为此,我们提出动态精准视觉思考框架(DaP-ICoT),包含两项核心设计:(1)动态视觉融合机制,根据推理需求自适应引入视觉信息,降低冗余并提升效率;(2)精准视觉引导机制,确保视觉表征语义连贯且上下文对齐。在多个基准和模型上的实验表明,DaP-ICoT达到当前最优性能,并将插入图像数量显著减少,令牌消耗降低72.6%,实现了更高效的ICoT推理。
原文摘要 · Abstract (English)
Recently, Interleaved-modal Chain-of-Thought (ICoT) reasoning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing attention. While achieving promising performance, current ICoT methods still suffer from two major limitations: (1) Static Visual Thought Positioning, which statically inserts visual information at fixed steps, resulting in inefficient and inflexible reasoning; and (2) Broken Visual Thought Representation, which involves discontinuous and semantically incoherent visual tokens. To address these limitations, we introduce Interleaved-modal Chain-of-Thought reasoning with Dynamic and Precise Visual Thoughts (DaP-ICoT), which incorporates two key components: (1) Dynamic Visual Thought Integration adaptively introduces visual inputs based on reasoning needs, reducing redundancy and improving efficiency. (2) Precise Visual Thought Guidance ensures visual semantically coherent and contextually aligned representations. Experiments across multiple benchmarks and models demonstrate that DaP-ICoT achieves state-of-the-art performance. In addition, DaP-ICoT significantly reduces the number of inserted images, leading to a 72.6% decrease in token consumption, enabling more efficient ICoT reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。