用强化学习提升视觉语言模型的符号化推理能力,效率提升75%。
Incentivizing Neuro-symbolic Language-based Reasoning in VLMs via Reinforcement Learning

- 基于神经符号语言设计强化学习框架,优化模型思维过程。
- 在多领域问答数据集上准确率提升3.33%,推理令牌减少75%。
- 适合关注高效推理与跨模态认知建模的研究者。
全球有7,407种语言,但是否存在人类未曾接触的外星语言?本文探索在神经符号语言框架下,视觉语言模型对视觉-语言概念的表征与推理能力,并研究分析性推理性能与效率的提升。以Qwen3-VL-2B-Instruct为基座模型,使用4×Nvidia H200 GPU节点进行训练,在包含数学、科学与常识问题的视觉语言评估数据集上,准确率提升3.33%,同时相比SymPy将推理所需令牌数减少75%。文中记录了计算挑战、可扩展性及未来在视觉语言模型中实现更高效“思维系统”的方向。完整训练与推理设置见:https://github.com/i-like-bfs-and-dfs/wolfram-reasoning。
原文摘要 · Abstract (English)
There are 7,407 languages in the world. But, what about the languages that are not there in the world? Are humans so narrow minded that we don't care about the languages aliens communicate in? Aliens are humans too! In the 2016 movie Arrival, Amy Adams plays a linguist, Dr. Louise Banks who, by learning to think in an alien language (Heptapod) formed of non-sequential sentences, gains the ability to transcend time and look into the future. In this work, I aim to explore the representation and reasoning of vision-language concepts in a neuro-symbolic language, and study improvement in analytical reasoning abilities and efficiency of "thinking systems". With Qwen3-VL-2B-Instruct as base model and 4 $\times$ Nvidia H200 GPU nodes, I achieve an accuracy improvement of 3.33\% on a vision-language evaluation dataset consisting of math, science, and general knowledge questions, while reducing the reasoning tokens by 75\% over SymPy. I've documented the compute challenges faced, scaling possibilities, and the future work to improve thinking in a neuro-symbolic language in vision-language models. The training and inference setup can be found here: https://github.com/i-like-bfs-and-dfs/wolfram-reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。