用人类与AI反馈强化学习,提升大模型解物理题能力。
Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback
- 结合人类与AI反馈,用强化学习优化大模型推理路径。
- 在PhyQA数据集上,推理得分达0.74,METEOR达58.67。
- 适用于需要复杂逻辑推理的物理教育场景研究。
大型语言模型在文本任务中表现优异,但在物理问题所需的复杂推理方面仍显不足,尤其在高级算术和概念理解上。尽管已有研究尝试通过提示工程和检索增强生成(RAG)改善模型在物理教育中的表现,但对推理能力局限性的系统性改进仍不足。本文提出一种基于人类与人工智能反馈的强化学习方法(RLHAIF),评估了近端策略优化(PPO)、直接偏好优化(DPO)和Remax优化等多种强化学习算法在PhyQA数据集上的表现,该数据集包含来自高中教材的挑战性物理题目。实验在LLaMA2和Mistral等主流大模型上进行,结果显示,采用MISTRAL-PPO的RLHAIF模型在推理能力和准确性上均有显著提升,取得58.67的METEOR分数和0.74的推理得分,为未来物理推理研究提供了有力范例。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated strong capabilities in text-based tasks but struggle with the complex reasoning required for physics problems, particularly in advanced arithmetic and conceptual understanding. While some research has explored ways to enhance LLMs in physics education using techniques such as prompt engineering and Retrieval Augmentation Generation (RAG), not enough effort has been made in addressing their limitations in physics reasoning. This paper presents a novel approach to improving LLM performance on physics questions using Reinforcement Learning with Human and Artificial Intelligence Feedback (RLHAIF). We evaluate several reinforcement learning methods, including Proximal Policy Optimization (PPO), Direct Preference Optimization (DPO), and Remax optimization. These methods are chosen to investigate RL policy performance with different settings on the PhyQA dataset, which includes challenging physics problems from high school textbooks. Our RLHAIF model, tested on leading LLMs like LLaMA2 and Mistral, achieved superior results, notably with the MISTRAL-PPO model, demonstrating marked improvements in reasoning and accuracy. It achieved high scores, with a 58.67 METEOR score and a 0.74 Reasoning score, making it a strong example for future physics reasoning research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。