arXiv:2509.06079cs.CLcs.CV2025-09被引 6

用图文联合推理提升科学问题解答能力,获ICML 2025挑战赛冠军

Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge

  • 引入图文辅助推理框架,打通视觉与文本模态
  • 在SeePhys挑战赛中排名第一,几何推理任务表现优异
  • 代码开源,适合科研与教育场景的多模态推理研究

多模态推理仍是人工智能的核心挑战。尽管文本推理取得显著进展,但即使最先进的模型如GPT-o3,在多模态场景中仍难以保持强劲表现。为此,我们提出一种基于图像描述的辅助推理框架,有效融合视觉与文本信息。该方法在ICML 2025 AI for Math Workshop & Challenge 2:SeePhys中获得第一名,验证了其有效性与鲁棒性。此外,我们在MathVerse几何推理基准上进一步验证了方法的泛化能力,展现了其广泛适用性。代码已公开于https://github.com/OpenDCAI/SciReasoner。

原文摘要 · Abstract (English)

Multimodal reasoning remains a fundamental challenge in artificial intelligence. Despite substantial advances in text-based reasoning, even state-of-the-art models such as GPT-o3 struggle to maintain strong performance in multimodal scenarios. To address this gap, we introduce a caption-assisted reasoning framework that effectively bridges visual and textual modalities. Our approach achieved 1st place in the ICML 2025 AI for Math Workshop \& Challenge 2: SeePhys, highlighting its effectiveness and robustness. Furthermore, we validate its generalization on the MathVerse benchmark for geometric reasoning, demonstrating the versatility of our method. Our code is publicly available at https://github.com/OpenDCAI/SciReasoner.

多模态推理科学计算图像理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。