让AI模型真正看懂图像,通过重建视觉思考过程提升多模态推理能力。
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
- 通过重建教师模型的视觉注意力轨迹来对齐学生模型的视觉理解。
- 在复杂推理任务上提升16.9%,30亿参数模型超越更大模型和GPT-4o。
- 适合追求高精度视觉推理、关注模型可解释性的研究者与开发者。
当前多模态潜在推理常依赖外部监督(如辅助图像),忽视内在视觉注意动态。本文发现知识蒸馏中的关键感知鸿沟:学生模型虽模仿教师的文字输出,却关注截然不同的视觉区域,实质依赖语言先验而非真实感知。为此,提出LaViT框架,对齐潜在视觉思维而非静态嵌入。该方法迫使学生在生成文本前,自回归重建教师的视觉语义与注意力轨迹,并采用课程感知门控机制防止捷径学习。大量实验表明,LaViT显著提升视觉定位能力,在复杂推理任务上最高提升16.9%,使3B参数模型超越更大的开源模型及GPT-4o等专有模型。
原文摘要 · Abstract (English)
Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dynamics. In this work, we identify a critical Perception Gap in distillation: student models frequently mimic a teacher's textual output while attending to fundamentally divergent visual regions, effectively relying on language priors rather than grounded perception. To bridge this, we propose LaViT, a framework that aligns latent visual thoughts rather than static embeddings. LaViT compels the student to autoregressively reconstruct the teacher's visual semantics and attention trajectories prior to text generation, employing a curriculum sensory gating mechanism to prevent shortcut learning. Extensive experiments show that LaViT significantly enhances visual grounding, achieving up to +16.9% gains on complex reasoning tasks and enabling a compact 3B model to outperform larger open-source variants and proprietary models like GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。