MAMUT可生成数学公式的等价变体,提升大模型对数学符号的理解能力。
MAMUT: A Novel Framework for Modifying Mathematical Formulas for the Generation of Specialized Datasets for Language Model Training
- 通过LaTeX公式变换生成等价但形式多样的数学表达
- 在数学检索任务上实现新SOTA性能,优于现有方法
- 适合需要增强数学推理能力的模型训练者使用
数学公式是科学领域中基础且广泛应用的表达方式,作为复杂概念与关系的通用语言。尽管当前先进的Transformer模型在自然语言处理方面表现优异,但在处理数学符号时仍面临挑战,因其结构复杂且表示多样。本研究聚焦于构建专用训练数据集以增强模型对数学内容的编码能力。我们提出Math Mutator(MAMUT)框架,能够生成给定数学公式的等价版本与伪造版本,有效捕捉同一概念在不同符号表达中的多样性。基于MAMUT,我们构建了四个包含多样化符号表示的大规模数学数据集。实验表明,使用这些数据集训练的模型在数学检索任务上达到新的SOTA性能。代码、生成数据集及预训练模型均已开源:https://github.com/aieng-lab/math-mutator。
原文摘要 · Abstract (English)
Mathematical formulas are a fundamental and widely used component in various scientific fields, serving as a universal language for expressing complex concepts and relationships. While state-of-the-art transformer models excel in processing and understanding natural language, they encounter challenges with mathematical notation, which involves a complex structure and diverse representations. This study focuses on the development of specialized training datasets to enhance the encoding of mathematical content. We introduce Math Mutator (MAMUT), a framework capable of generating equivalent and falsified versions of a given mathematical formula in LaTeX notation, effectively capturing the mathematical variety in notation of the same concept. Based on MAMUT, we have generated four large mathematical datasets containing diverse notation. Experiments show that models trained on these datasets exhibit new SoTA performance on mathematical retrieval tasks. We publish our code, generated datasets, and pretrained mathematical models: https://github.com/aieng-lab/math-mutator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。