arXiv:2511.17012cs.CLcs.AI2025-11

用指令微调提升大模型对湖南历史名人知识图谱的构建能力

Supervised Fine Tuning of Large Language Models for Domain Specific Knowledge Graph Construction:A Case Study on Hunan's Historical Celebrities

  • 设计领域专属指令模板,构建湖南名人知识提取数据集
  • 微调后模型准确率达89.39,Qwen3-8B表现最佳
  • 为地方文化数字化提供低成本高效率的技术路径

大语言模型与知识图谱在历史文化遗产研究中具有巨大潜力,可支持文化信息的抽取、分析与解读。以湖湘文化孕育的湖南近现代历史名人为例,预训练大模型能高效从文本中提取人物生平、事件与社会关系等关键信息,并构建结构化知识图谱。然而,湖南历史名人领域的系统性数据资源仍较匮乏,通用模型在此类低资源场景下知识抽取与结构化输出性能不足。为此,本研究提出一种监督微调方法,以增强领域特定信息抽取能力:首先,设计细粒度、模式引导的指令模板,构建指令微调数据集,缓解领域专用语料短缺问题;其次,对四个公开大模型(Qwen2.5-7B、Qwen3-8B、DeepSeek-R1-Distill-Qwen-7B、Llama-3.1-8B-Instruct)采用参数高效指令微调,并建立评估标准。实验表明,所有模型经微调后性能显著提升,其中Qwen3-8B在100样本、50轮训练下达到89.3866的得分,表现最优。该研究为垂直领域大模型微调提供了新思路,展示了其在区域历史文化知识提取与知识图谱构建中的低成本应用前景。

原文摘要 · Abstract (English)

Large language models and knowledge graphs offer strong potential for advancing research on historical culture by supporting the extraction, analysis, and interpretation of cultural heritage. Using Hunan's modern historical celebrities shaped by Huxiang culture as a case study, pre-trained large models can help researchers efficiently extract key information, including biographical attributes, life events, and social relationships, from textual sources and construct structured knowledge graphs. However, systematic data resources for Hunan's historical celebrities remain limited, and general-purpose models often underperform in domain knowledge extraction and structured output generation in such low-resource settings. To address these issues, this study proposes a supervised fine-tuning approach for enhancing domain-specific information extraction. First, we design a fine-grained, schema-guided instruction template tailored to the Hunan historical celebrities domain and build an instruction-tuning dataset to mitigate the lack of domain-specific training corpora. Second, we apply parameter-efficient instruction fine-tuning to four publicly available large language models - Qwen2.5-7B, Qwen3-8B, DeepSeek-R1-Distill-Qwen-7B, and Llama-3.1-8B-Instruct - and develop evaluation criteria for assessing their extraction performance. Experimental results show that all models exhibit substantial performance gains after fine-tuning. Among them, Qwen3-8B achieves the strongest results, reaching a score of 89.3866 with 100 samples and 50 training iterations. This study provides new insights into fine-tuning vertical large language models for regional historical and cultural domains and highlights their potential for cost-effective applications in cultural heritage knowledge extraction and knowledge graph construction.

知识图谱大模型微调文化数字化湖湘文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。