针对韩国法律场景,构建了高精度专用大模型LegalMidm
LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
- 基于真实法律场景需求设计数据集与训练流程
- 通过专业人士协作确保数据准确性和实用性
- 适合法律科技、智能辅助判决等专业领域使用
近年来,开源大语言模型的快速发展推动了通用模型向领域专用模型的转化。然而,许多领域专用模型所采用的数据集和训练方法未能匹配实际应用的细微需求。在对精确性和可靠性要求极高的法律领域,这种脱节限制了其实际应用价值。本研究提出一种以实际应用需求为导向的系统性训练框架,聚焦韩国法律领域,构建了韩国法律专用大模型LegalMidm。我们提出一种高质量、用例驱动的法律数据集构建方法及优化的训练流程,强调与法律专业人士的协作以及严格的语料筛选,以确保内容的相关性与事实准确性,并在多个关键法律任务中验证了其有效性。
原文摘要 · Abstract (English)
In recent years, the rapid proliferation of open-source large language models (LLMs) has spurred efforts to turn general-purpose models into domain specialists. However, many domain-specialized LLMs are developed using datasets and training protocols that are not aligned with the nuanced requirements of real-world applications. In the legal domain, where precision and reliability are essential, this lack of consideration limits practical utility. In this study, we propose a systematic training framework grounded in the practical needs of the legal domain, with a focus on Korean law. We introduce LegalMidm, a Korean legal-domain LLM, and present a methodology for constructing high-quality, use-case-driven legal datasets and optimized training pipelines. Our approach emphasizes collaboration with legal professionals and rigorous data curation to ensure relevance and factual accuracy, and demonstrates effectiveness in key legal tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。