用视觉工具增强大模型,精准验证科学论文中的图文论断
ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

- 为多模态科学论断验证设计三类专用视觉工具,精准提取图表证据
- 在五个主流视觉语言模型上超越基线,提升验证准确率与工具使用效率
- 适合需要高精度科学推理的科研人员和论文审核者使用
多模态科学论断验证(MSCV)要求模型基于论文中的图表、表格、图示和文本上下文来验证科学论断。现有方法常因难以定位关键视觉证据、准确解读结构化科学图表以及整合多模态信息进行可靠推理而失败。我们提出ToolSciVer,据我们所知首个工具增强型MSCV框架。该框架为视觉语言模型(VLM)配备三类类型感知的视觉工具:表格行列聚焦、图表结构解析与高分辨率区域缩放,将密集科学图表转化为明确的、针对论断的证据。采用组相对策略优化(GRPO)训练策略,在包含答案正确性、格式有效性、长度控制、工具使用效率及工具有效性惩罚的复合奖励下进行优化。在SciVer与MuSciClaims数据集上对来自Qwen、InternVL、Gemma三个模型家族的五种VLM进行实验,结果表明本方法显著优于四种对比基线,包括基于提示和强化学习的工具使用方法,验证了学习型、类型感知工具使用在科学论断验证中的有效性。
原文摘要 · Abstract (English)
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they struggle to locate decisive visual evidence, accurately read structured scientific visuals, and integrate multimodal observations into reliable reasoning. We introduce ToolSciVer, the first tool-augmented framework for MSCV to our knowledge. ToolSciVer equips a VLM with three type-aware visual tools, table row/column focus, chart-to-structure parsing, and high-resolution region zoom, which convert dense scientific visuals into explicit, claim-facing evidence, and trains the policy with Group Relative Policy Optimization (GRPO) under a composite reward of answer correctness, format validity, length control, tool-use efficiency, and tool-validity penalties. Experiments on SciVer and MuSciClaims datasets on five VLMs from three model families (Qwen, InternVL, Gemma) demonstrate that our method achieves superior performance compared to four competitive baselines including prompting-based and RL-based tool-use methods, highlighting the effectiveness of learned, type-aware tool use for scientific claim verification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。