用小规模语料微调大模型,实现芬兰语医疗对话的高精度转录。
Evaluating Fine-Tuned LLM Model For Medical Transcription With Small Low-Resource Languages Validated Dataset
- 在小规模芬兰语临床对话数据上微调LLaMA 3.1-8B模型。
- BLEU=0.1214,ROUGE-L=0.4982,BERTScore F1=0.8230,语义匹配度高。
- 为低资源语言医疗文本生成提供可行路径,适合医疗AI研究者。
临床文档是患者安全、诊断与连续照护的关键因素,但电子病历(EHR)带来的行政负担是导致医生职业倦怠的重要原因。这一问题在低资源语言中尤为突出,例如芬兰语。本研究旨在通过在由赫尔辛基应用科学大学学生模拟的临床对话小规模验证语料上微调LLaMA 3.1-8B,评估领域对齐自然语言处理大语言模型在芬兰语医疗转录中的有效性。采用受控预处理与优化方法进行微调,并通过七折交叉验证评估性能。结果表明,微调后模型的BLEU=0.1214,ROUGE-L=0.4982,BERTScore F1=0.8230,虽词元重叠度低,但与参考转录具有强语义相似性。研究证实,微调是实现芬兰语口语医疗话语转录的有效方法,支持在隐私保护前提下构建特定领域的大型语言模型用于芬兰语临床文档生成,并为未来工作提供方向。
原文摘要 · Abstract (English)
Clinical documentation is a critical factor for patient safety, diagnosis, and continuity of care. The administrative burden of EHRs is a significant factor in physician burnout. This is a critical issue for low-resource languages, including Finnish. This study aims to investigate the effectiveness of a domain-aligned natural language processing (NLP); large language model for medical transcription in Finnish by fine-tuning LLaMA 3.1-8B on a small validated corpus of simulated clinical conversations by students at Metropolia University of Applied Sciences. The fine-tuning process for medical transcription used a controlled preprocessing and optimization approach. The fine-tuning effectiveness was evaluated by sevenfold cross-validation. The evaluation metrics for fine-tuned LLaMA 3.1-8B were BLEU = 0.1214, ROUGE-L = 0.4982, and BERTScore F1 = 0.8230. The results showed a low n-gram overlap but a strong semantic similarity with reference transcripts. This study indicate that fine-tuning can be an effective approach for translation of medical discourse in spoken Finnish and support the feasibility of fine-tuning a privacy-oriented domain-specific large language model for clinical documentation in Finnish. Beside that provide directions for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。