小模型微调提升量子场论推理能力,数据与方法开源。
Fine-Tuning Small Reasoning Models for Quantum Field Theory

- 用自动生成与改编数据训练70亿参数小模型
- 微调后模型在量子场论任务上表现显著提升
- 适合对物理推理与小模型研究者参考
尽管大语言模型在理论物理中应用日益广泛,但关于领域特定物理推理能力如何随训练发展仍缺乏学术研究。为此,我们首次针对70亿参数的小型推理模型开展专门面向理论物理的微调研究。由于开源可验证的训练数据稀缺,我们构建了稳健的数据生成流程,既能生成合成问题,也能使现有真人撰写的题目适用于模型训练。以量子场论(QFT)为主要研究领域,我们生成了超过2500个合成问题,并整理了来自arXiv和标准教学资源的人工改编问题。我们进行了强化学习(RL)与监督微调(SFT)实验,评估性能提升及向其他物理领域的泛化能力。通过对比微调前后模型的思维链,深入分析了推理错误在训练过程中的演变。最后,我们公开发布数据生成管道、可验证的QFT训练数据,以及约2亿条QFT推理轨迹。
原文摘要 · Abstract (English)
Despite the growing application of Large Language Models (LLMs) to theoretical physics, there is little academic exploration into how domain-specific physics reasoning ability develops while training these models. To investigate this, we perform the first academic fine-tuning study of small (7B-parameter) reasoning models dedicated specifically to theoretical physics. Because open-source verifiable training data required to train such capabilities is scarce, we developed a robust data generation pipeline that can both create synthetic problems and make existing human-authored problems suitable for model training. Selecting Quantum Field Theory (QFT) as our primary domain, we generated over 2,500 synthetic problems alongside a curated collection of human-adapted problems sourced from arXiv and standard pedagogical resources. We conduct both Reinforcement Learning (RL) and Supervised Fine-Tuning (SFT) experiments, benchmarking performance gains as well as generalization to other physics domains. We perform an extensive analysis of model chains-of-though before and after fine-tuning, to understand how reasoning errors evolve during RL and SFT. Finally, we publicly release our data pipeline, verifiable QFT training data, and $\sim$200M tokens of QFT reasoning traces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。