将图像转为文本形式,让大模型更精准地进行跨模态推理。
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

- 把图像转化为形式化文本,实现视觉信息的语言化推理。
- 在多个复杂任务上超越GPT-4o和Qwen2.5-VL等先进模型。
- 构建了覆盖中学到大学水平的跨模态推理评测基准。
大型语言模型在复杂文本任务中展现出强大推理能力,但跨模态推理——即整合视觉与文本信息——仍是重大挑战。现有视觉语言模型难以有效分析视觉内容,导致在复杂推理任务中表现不佳。此外,缺乏全面的评估基准也阻碍了对跨模态推理能力的准确衡量。本文提出R1-Onevision,一种旨在弥合视觉感知与深度推理之间差距的多模态推理模型。我们设计了一种跨模态推理流程,将图像转换为形式化文本表示,从而支持精确的语言推理。基于该流程,我们构建了R1-Onevision数据集,涵盖多个领域的详细、分步的多模态推理标注。通过监督微调和强化学习,我们进一步训练了R1-Onevision模型,以增强其推理能力和鲁棒泛化性能。为全面评估不同层次的多模态推理表现,我们引入R1-Onevision-Bench,一个与人类教育阶段对齐的基准,涵盖从初中到大学及以上的考试题目。实验结果表明,R1-Onevision在多个具有挑战性的多模态推理基准上达到领先水平,优于GPT-4o和Qwen2.5-VL等模型。
原文摘要 · Abstract (English)
Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing visual-language models often struggle to effectively analyze and reason visual content, resulting in suboptimal performance on complex reasoning tasks. Moreover, the absence of comprehensive benchmarks hinders the accurate assessment of multimodal reasoning capabilities. In this paper, we introduce R1-Onevision, a multimodal reasoning model designed to bridge the gap between visual perception and deep reasoning. To achieve this, we propose a cross-modal reasoning pipeline that transforms images into formal textural representations, enabling precise language-based reasoning. Leveraging this pipeline, we construct the R1-Onevision dataset which provides detailed, step-by-step multimodal reasoning annotations across diverse domains. We further develop the R1-Onevision model through supervised fine-tuning and reinforcement learning to cultivate advanced reasoning and robust generalization abilities. To comprehensively evaluate multimodal reasoning performance across different grades, we introduce R1-Onevision-Bench, a benchmark aligned with human educational stages, covering exams from junior high school to university and beyond. Experimental results show that R1-Onevision achieves state-of-the-art performance, outperforming models such as GPT-4o and Qwen2.5-VL on multiple challenging multimodal reasoning benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。