提升气候问答大模型的忠实度,让回答更依赖检索文本。
Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
- 通过筛选训练数据,优化指令微调过程以增强模型忠实性。
- 改进后模型在原子主张上的忠实度从30%提升至57%。
- 适合关注气候科学问答可靠性的研究者与政策制定者。
使用检索增强生成的大语言模型有望通过使长篇技术性气候文件更易获取,为研究人员、决策者和公众释放宝贵知识。尽管该方法可通过依赖检索段落缓解事实幻觉,但其效果取决于模型输出是否忠实于这些段落。为此,我们探索了在此场景下自动评估模型忠实性的方法,并聚焦于专攻气候科学的ClimateGPT模型,分析其指令微调中影响忠实性的因素。通过剔除不忠实的训练数据子集,我们构建了ClimateGPT Faithful+,在支持的原子主张上,其忠实度自动指标从30%提升至57%。
原文摘要 · Abstract (English)
Large language models that use retrieval augmented generation have the potential to unlock valuable knowledge for researchers, policymakers, and the public by making long and technical climate-related documents more accessible. While this approach can help alleviate factual hallucinations by relying on retrieved passages as additional context, its effectiveness depends on whether the model's output remains faithful to these passages. To address this, we explore the automatic assessment of faithfulness of different models in this setting. We then focus on ClimateGPT, a large language model specialised in climate science, to examine which factors in its instruction fine-tuning impact the model's faithfulness. By excluding unfaithful subsets of the model's training data, we develop ClimateGPT Faithful+, which achieves an improvement in faithfulness from 30% to 57% in supported atomic claims according to our automatic metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。