让视觉隐状态自发生成关键语义,提升多模态模型理解力
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
- 双路径设计分别处理完整与损坏图像,通过注意力对齐引导隐状态优化
- 在多个感知密集型基准上超越现有模型,细粒度视觉理解显著提升
- 无需额外标注或模块,可直接增强模型推理能力,适合视觉理解任务
多模态大语言模型(MLLM)通过融合强大语言主干与大规模视觉编码器,在多项任务中表现卓越。其中,隐式思维链(latent CoT)方法能在连续隐藏状态中实现隐式推理,促进视觉-语言无缝融合并加速推理。然而,现有方法依赖启发式预设的监督信号,难以有效保留中间隐状态中的关键视觉信息。为此,我们提出CrystaL(Crystallized Latent Reasoning),一种单阶段框架,包含两条路径:分别处理完整图像与受损图像。通过显式对齐两条路径的注意力模式与预测分布,CrystaL将隐表示“结晶”为与任务相关的视觉语义,无需辅助标注或外部模块。在多个感知密集型基准上的实验表明,CrystaL持续优于当前最优基线,在细粒度视觉理解方面取得显著提升,同时保持强推理能力。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have achieved remarkable performance by integrating powerful language backbones with large-scale visual encoders. Among these, latent Chain-of-Thought (CoT) methods enable implicit reasoning in continuous hidden states, facilitating seamless vision-language integration and faster inference. However, existing heuristically predefined supervision signals in latent CoT provide limited guidance for preserving critical visual information in intermediate latent states. To address this limitation, we propose CrystaL (Crystallized Latent Reasoning), a single-stage framework with two paths to process intact and corrupted images, respectively. By explicitly aligning the attention patterns and prediction distributions across the two paths, CrystaL crystallizes latent representations into task-relevant visual semantics, without relying on auxiliary annotations or external modules. Extensive experiments on perception-intensive benchmarks demonstrate that CrystaL consistently outperforms state-of-the-art baselines, achieving substantial gains in fine-grained visual understanding while maintaining robust reasoning capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。