arXiv:2606.16003cs.AI2026-06ACL

测试大模型从科学文本生成方程的能力,发现其语义准确率不高。

SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity

论文配图:SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity
图 1 · 摘自论文原文
  • 构建科学论文数据集,含上下文、方程与变量说明
  • 大模型在词法和语法上表现尚可,但语义准确性差
  • 人工评估与模型自评对齐度低,提示评估方法需改进

本研究探究大语言模型(LLMs)从科学文本生成数学方程的能力。先前工作面临非结构化对齐、多方程依赖关系及人类对齐评估的挑战。为此,我们构建了一个由人工智能研究论文组成的数据库,将上下文段落与真实方程及变量描述配对。开发了可解释的方程生成流程,并在多种开源与闭源模型上进行评估。提出结合自动指标、基于LLM的评分标准和人工判断的评估协议,以衡量准确性、可解释性及人-模型一致性。结果表明,大模型在词汇与句法相似性方面表现中等,但在语义准确性上表现不佳。对比显示,基于LLM的评估与人工判断一致性有限,凸显利用大模型评估方程质量的挑战。这些发现为改进方程生成模型和建立更可靠的科学文本评估方法提供了参考。代码与数据已公开以支持复现。

原文摘要 · Abstract (English)

This work investigates the ability of large language models (LLMs) to generate mathematical equations from scientific texts. Prior work faces challenges in unstructured grounding, multi-equation dependency, and humanaligned evaluation. To this end, we construct a dataset of AI research papers, pairing contextual passages with ground-truth equations and variable descriptions. We develop an explainable equation generation workflow and evaluate it across diverse open- and closed-source LLM backbones. We introduce an evaluation protocol combining automatic metrics, LLM-based rubrics, and human judgments to assess accuracy, explainability, and human-LLM alignment. Results indicate that LLMs perform moderately on lexical- and syntactic-based similarity, while struggling with semantic accuracy. Comparisons between LLM-based evaluations and human judgments reveal limited alignment, highlighting challenges in using LLMs to assess equation quality. These findings offer insights for improving equation generation models and developing more reliable evaluation methods for scientific text. We provide code and data for reproducibility.

方程生成大模型评估科学智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。