arXiv:2603.03975cs.AIcs.CV2026-03被引 6

15B参数的多模态模型,小而强,专精数学推理与界面理解。

Phi-4-reasoning-vision-15B Technical Report

  • 用高质量数据筛选和合成增强,提升小模型性能。
  • 在数学推理和界面理解上超越大模型,训练量减少60%以上。
  • 单模型支持直答与思维链推理,适合复杂任务与快速响应。

我们提出Phi-4-reasoning-vision-15B,一个轻量级开源多模态推理模型,旨在为构建高效小型多模态模型提供实践洞见。目标是通过精心设计的架构和严格的语料筛选,使小型开放权重模型在显著降低训练与推理计算成本的前提下,达到与更大模型相当的性能。关键改进来自系统性数据过滤、错误修正与合成增强,凸显数据质量是提升模型表现的核心杠杆。消融实验表明,高分辨率、动态分辨率编码器持续带来性能提升,因准确感知是高质量推理的前提。此外,通过引入显式模式标记的混合数据策略,单一模型可同时实现简单任务的快速直接回答与复杂问题的链式思维推理。

原文摘要 · Abstract (English)

We present Phi-4-reasoning-vision-15B, a compact open-weight multimodal reasoning model, and share the motivations, design choices, experiments, and learnings that informed its development. Our goal is to contribute practical insight to the research community on building smaller, efficient multimodal reasoning models and to share the result of these learnings as an open-weight model that is good at common vision and language tasks and excels at scientific and mathematical reasoning and understanding user interfaces. Our contributions include demonstrating that careful architecture choices and rigorous data curation enable smaller, open-weight multimodal models to achieve competitive performance with significantly less training and inference-time compute and tokens. The most substantial improvements come from systematic filtering, error correction, and synthetic augmentation -- reinforcing that data quality remains the primary lever for model performance. Systematic ablations show that high-resolution, dynamic-resolution encoders yield consistent improvements, as accurate perception is a prerequisite for high-quality reasoning. Finally, a hybrid mix of reasoning and non-reasoning data with explicit mode tokens allows a single model to deliver fast direct answers for simpler tasks and chain-of-thought reasoning for complex problems.

多模态推理小模型数学推理数据质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。