用e图生成数学表达式数据集,提升符号表达的向量表示效果
E-Gen: Leveraging E-Graphs to Improve Continuous Representations of Symbolic Expressions
- 基于e图自动生成大规模多样数学表达式数据集
- 在数学语言任务中超越现有方法,部分任务优于大模型
- 适合需要精准数学表达理解的研究者和开发者
向量表示在自然语言处理中至关重要,以往研究聚焦于使用数学等价形式对数学表达式进行嵌入。尽管有效,这些方法受限于训练数据规模与多样性。本文提出E-Gen,一种基于e图的数据集生成方案,可合成大规模且多样化的数学表达式数据集,其规模与算子种类均超过先前方法。利用该数据集,我们采用两种策略训练嵌入模型:(1) 生成数学等价表达式;(2) 通过对比学习显式聚合同等表达式。我们在分布内与分布外的数学语言处理任务上评估嵌入效果,并与已有方法对比。结果表明,我们的嵌入方法在多个任务上优于当前最先进大语言模型,凸显了针对数学数据模态优化嵌入方法的必要性。源代码与数据集已开源。
原文摘要 · Abstract (English)
Vector representations have been pivotal in advancing natural language processing (NLP), with prior research focusing on embedding techniques for mathematical expressions using mathematically equivalent formulations. While effective, these approaches are constrained by the size and diversity of training data. In this work, we address these limitations by introducing E-Gen, a novel e-graph-based dataset generation scheme that synthesizes large and diverse mathematical expression datasets, surpassing prior methods in size and operator variety. Leveraging this dataset, we train embedding models using two strategies: (1) generating mathematically equivalent expressions, and (2) contrastive learning to explicitly group equivalent expressions. We evaluate these embeddings on both in-distribution and out-of-distribution mathematical language processing tasks, comparing them against prior methods. Finally, we demonstrate that our embedding-based approach outperforms state-of-the-art large language models (LLMs) on several tasks, underscoring the necessity of optimizing embedding methods for the mathematical data modality. The source code and datasets are available at https://github.com/MLPgroup/E-Gen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。