用符号化推理提升大模型数学能力,过程可验证更可信
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
- 将自然语言转为符号表达,再进行逻辑推导
- 在7个基准中6个超越传统思维链,最高提效4.58%
- 无需外部工具,适合需要透明推理的数学任务
大语言模型在数学推理上仍面临挑战,尽管提示技术如思维链(CoT)有所进展。本文提出链式数学注释思维(CoMAT),通过两个阶段增强推理:符号转换(将自然语言查询转化为符号形式)和推理执行(从符号表示中推导答案)。CoMAT完全依赖单一LLM,无需外部求解器。在四个不同LLM上测试,CoMAT在七个基准中的六个优于传统CoT,MMLU-Redux(MATH)提升4.48%,高考选择题(GaoKao MCQ)提升4.58%。除性能提升外,CoMAT还保证了推理过程的忠实性与可验证性,为复杂数学任务提供透明推理路径。
原文摘要 · Abstract (English)
Mathematical reasoning remains a significant challenge for large language models (LLMs), despite progress in prompting techniques such as Chain-of-Thought (CoT). We present **Chain of Mathematically Annotated Thought (CoMAT)**, which enhances reasoning through two stages: *Symbolic Conversion* (converting natural language queries into symbolic form) and *Reasoning Execution* (deriving answers from symbolic representations). CoMAT operates entirely with a single LLM and without external solvers. Across four LLMs, CoMAT outperforms traditional CoT on six out of seven benchmarks, achieving gains of 4.48% on MMLU-Redux (MATH) and 4.58% on GaoKao MCQ. In addition to improved performance, CoMAT ensures faithfulness and verifiability, offering a transparent reasoning process for complex mathematical tasks
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。