arXiv:2412.05479cs.CV2024-12EMNLP被引 7

让视觉模型专精看图,语言模型专注推理,提升复杂问题回答能力

LATTE: Learning to Think with Vision Specialists

  • 将视觉感知交给顶尖视觉模型,语言模型只负责推理
  • 在6个基准上比基线提升4-5%准确率
  • 适合需要强推理能力的多模态问答任务

开源视觉语言模型在简单问答中表现良好,但在需结合感知与推理的复杂问题上仍表现不佳。我们提出LATTE,一种让视觉语言模型学会‘思考’的新型架构。通过将感知任务交由先进的视觉模型完成,语言模型可专注于高质量感知输出的推理。为训练LATTE,我们构建并筛选了包含293,000条多模态推理轨迹的数据集,这些轨迹基于视觉专家模型的输出。在6个涵盖感知与推理能力的基准测试中,LATTE相比基线模型取得4-5%的显著提升。消融实验表明,多模态推理轨迹的效果取决于数据来源、格式和思维质量。

原文摘要 · Abstract (English)

While open-source vision-language models perform well on simple question-answering, they still struggle with complex questions that require both perceptual and reasoning capabilities. We propose LATTE, a family of vision-language models that have LeArned to Think wiTh vision spEcialists. By offloading perception to state-of-the-art vision models, our approach enables vision-language models to focus solely on reasoning over high-quality perceptual information. To train LATTE, we synthesize and filter a large dataset of 293K multi-modal reasoning traces over perceptual outputs of vision specialists. LATTE trained on this data achieves significant 4-5% gains over baselines across 6 benchmarks covering both perception and reasoning abilities. Ablation studies reveal that the effectiveness of multi-modal reasoning traces depends on the data sources, formats, and quality of thoughts.

视觉语言模型多模态推理思考机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。