用符号与神经网络结合生成几何推理数据,提升多模态模型能力
NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
- 基于实体-属性-关系语言构建几何符号空间,实现结构化数据生成
- 10万样本数据集支持,4k样本微调即提效超15%,4B模型超越8B
- 适合研究几何推理、多模态模型训练的学者和工程师
获取大规模高质量的推理数据对提升多模态大模型(MLLMs)的几何推理能力至关重要。现有数据生成方法受限于预设模板或符号证明器,存在多样性差和数值泛化能力弱的问题。为此,我们提出一种新的神经符号框架NeSyGeo,通过面向平面几何的实体-属性-关系范式语言,定义符号空间中的生成动作。设计符号-视觉-文本流水线,将符号序列映射为视觉与文本表示,并通过逆向搜索与正向验证生成推理路径。基于此框架,构建了包含10万样本的NeSyGeo CoT与NeSyGeo-Caption数据集,并发布新基准NeSyGeo-Test用于评估MLLM几何推理能力。实验表明,该方法在强化学习与监督微调下均显著提升多个MLLM性能:仅需4k样本和两轮强化微调,基础模型在MathVision、MathVerse、GeoQA上分别提升+15.8%、+8.4%、+7.3%。值得注意的是,4B模型经优化后可超越同系列8B模型。
原文摘要 · Abstract (English)
Obtaining large-scale, high-quality reasoning data is crucial for improving the geometric reasoning capabilities of multi-modal large language models (MLLMs). However, existing data generation methods, whether based on predefined tem plates or constrained symbolic provers, inevitably face diversity and numerical generalization limitations. To address these limitations, we propose NeSyGeo, a novel neuro-symbolic framework for generating geometric reasoning data. First, we propose a domain-specific language grounded in the entity-attributes-relations paradigm to comprehensively represent all components of plane geometry, along with generative actions defined within this symbolic space. We then design a symbolic-visual-text pipeline that synthesizes symbolic sequences, maps them to visual and textual representations and generates reasoning path with reverse search and forward validation. Based on this framework, we construct NeSyGeo CoT and NeSyGeo-Caption datasets, containing 100k samples, and release a new benchmark NeSyGeo-Test for evaluating geometric reasoning abilities in MLLMs. Experiments demonstrate that the proposal significantly and consistently improves the performance of multiple MLLMs under both reinforcement and supervised fine-tuning. With only 4k samples and two epochs of reinforcement fine-tuning, base models achieve improvements of up to +15.8% on MathVision, +8.4% on MathVerse, and +7.3% on GeoQA. Notably, a 4B model can be improved to outperform an 8B model from the same series on geometric reasoning tasks.s
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。