arXiv:2501.17840cs.CLcs.LG2025-01被引 5

用LoRA持续预训练,让大模型在医金融领域学得更深

Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?

  • 用LoRA在医学金融数据上持续预训练,提升模型深层理解能力
  • 删减冗余信息后,模型对三类知识的掌握提升显著
  • 适合想提升领域模型深度理解的研究者和开发者

大语言模型在各类任务中表现优异,但其从领域特定数据中提取并内化深层洞察的能力仍待深入探索。本研究探究持续预训练如何增强模型在三种不同形式洞察(陈述性、统计性、概率性)上的学习能力。聚焦医疗与金融两个关键领域,使用LoRA在两个现有数据集上训练模型。为评估各类洞察,构建基准测试以衡量持续预训练帮助模型超越表层知识的效果。同时考察文档修改对捕捉洞察的影响。结果显示,仅在原始文档上进行持续预训练效果有限;而通过删减冗余内容保留核心信息,能显著提升模型的洞察学习能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable performance on various tasks, yet their ability to extract and internalize deeper insights from domain-specific datasets remains underexplored. In this study, we investigate how continual pre-training can enhance LLMs' capacity for insight learning across three distinct forms: declarative, statistical, and probabilistic insights. Focusing on two critical domains: medicine and finance, we employ LoRA to train LLMs on two existing datasets. To evaluate each insight type, we create benchmarks to measure how well continual pre-training helps models go beyond surface-level knowledge. We also assess the impact of document modification on capturing insights. The results show that, while continual pre-training on original documents has a marginal effect, modifying documents to retain only essential information significantly enhances the insight-learning capabilities of LLMs.

大模型持续学习领域适应LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。