用数据库反馈强化学习,让大模型更好理解化学分子数据。
RLDBF: Enhancing LLMs Via Reinforcement Learning With DataBase FeedBack
- 通过强化学习引入数据库反馈,提升模型对结构化科学数据的利用能力。
- 在未见过的数据上表现优异,跨任务泛化能力显著增强。
- 适合希望融合科学数据库与大模型的AI科研人员使用。
当前大语言模型虽在海量非结构化文本训练下展现强大语言能力,但难以有效利用蕴含数百年科学积累的结构化数据(如化学分子属性数据库)。这些数据对推进科学智能至关重要,但现有方法仅将其作为非结构化文本的辅助补充。本研究以化学分子科学为测试场景,系统探索结构化科学数据对大模型的增强作用,考察其在持续预训练、监督微调和强化学习等不同训练阶段的影响。针对大模型固有的数值敏感性不足问题,提出创新方法「基于数据库反馈的强化学习」(RLDBF)。实验表明,该方法显著提升了模型在未见数据及其他化学任务上的泛化能力,验证了其在结构化科学数据处理中的潜力。
原文摘要 · Abstract (English)
While current large language models (LLMs) demonstrate remarkable linguistic capabilities through training on massive unstructured text corpora, they remain inadequate in leveraging structured scientific data (e.g., chemical molecular properties in databases) that encapsulate centuries of accumulated scientific expertise. These structured datasets hold strategic significance for advancing AI for Science yet current approaches merely treat them as auxiliary supplements to unstructured text. This study pioneers a systematic investigation into enhancing LLMs with structured scientific data, using chemical molecular science as a testbed. We investigate the impact of incorporating molecular property data on LLM across distinct training phases, including continual pre-training, supervised fine-tuning, and reinforcement learning. Notably, to address the inherent limitation of numerical insensitivity in large models, we propose an innovative methodology termed "Reinforcement Learning with Database Feedback" (RLDBF). Experimental evaluations demonstrate the efficacy of the proposed approach, with the model exhibiting remarkable generalization capabilities on previously unseen data and other chemical tasks. The results substantiate the potential of our method in advancing the field of structured scientific data processing within LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。