arXiv:2411.06445cs.CLcs.AI2024-11

用低资源微调减少大模型科学文本幻觉,提升可复现性

Prompt-Efficient Fine-Tuning for GPT-like Deep Models to Reduce Hallucination and to Improve Reproducibility in Scientific Text Generation Using Stochastic Optimisation Techniques

  • 基于LoRA对GPT-2进行轻量微调,专用于质谱文献生成
  • 在质谱数据集上,输出一致性与可复现性显著提升
  • 适合需要高可靠性的科研文本生成场景

大型语言模型在科学文本生成中应用日益广泛,但常面临准确性不足、一致性差和幻觉控制难的问题。本文针对GPT类模型提出一种参数高效微调(PEFT)方法,旨在减轻幻觉并提升可复现性,尤其聚焦质谱计算领域。通过使用质谱文献专有语料库,采用低秩适应(LoRA)适配器对GPT-2进行微调,构建了MS-GPT模型。借助BLEU、ROUGE及困惑度等评估指标,结合威尔科克森秩和检验的统计分析,结果显示微调后的MS-GPT在文本连贯性和可复现性上优于基线GPT-2。此外,提出基于提示控制下输出余弦相似度的可复现性度量,验证了模型稳定性。研究表明,该方法可在降低计算成本的同时提升模型可靠性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly adopted for complex scientific text generation tasks, yet they often suffer from limitations in accuracy, consistency, and hallucination control. This thesis introduces a Parameter-Efficient Fine-Tuning (PEFT) approach tailored for GPT-like models, aiming to mitigate hallucinations and enhance reproducibility, particularly in the computational domain of mass spectrometry. We implemented Low-Rank Adaptation (LoRA) adapters to refine GPT-2, termed MS-GPT, using a specialized corpus of mass spectrometry literature. Through novel evaluation methods applied to LLMs, including BLEU, ROUGE, and Perplexity scores, the fine-tuned MS-GPT model demonstrated superior text coherence and reproducibility compared to the baseline GPT-2, confirmed through statistical analysis with the Wilcoxon rank-sum test. Further, we propose a reproducibility metric based on cosine similarity of model outputs under controlled prompts, showcasing MS-GPT's enhanced stability. This research highlights PEFT's potential to optimize LLMs for scientific contexts, reducing computational costs while improving model reliability.

科学生成幻觉抑制参数高效微调可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。