arXiv:2608.00932cs.CL2026-08

打造轻量级波斯语医学大模型,助力本地化医疗AI落地

Gaokerena: A Small Persian Medical Language Model Family

论文配图:Gaokerena: A Small Persian Medical Language Model Family
图 1 · 摘自论文原文
  • 基于9000万词波斯语医学语料与2万条医生问答对训练
  • 推理模型在医疗MMLU测试中达52.98%准确率
  • 支持内部状态预测回答可信度,适合资源受限场景

人工智能在医疗问答系统中的应用快速发展,但研究仍主要集中在英语,低资源语言如波斯语严重被忽视。本文提出Gaokerena系列轻量级波斯语医学语言模型,专为消费级硬件部署优化。首先,Gaokerena-V基于新构建的9000万词波斯语医学语料和20,000条专家审核的医生问答对训练,使翻译后的医疗MMLU基准得分从46.28%提升至49.31%。其次,为增强临床推理能力,开发了结合思维链与两种新型基于AI反馈的强化学习(RLAIF)框架的Gaokerena-R。尽管使用相同基线架构且数据集更小,其基准得分达到52.98%,表现更优。此外,两个模型均配备自研不确定性头,仅通过内部隐藏状态即可预测回答置信度。当前性能尚不足以直接用于临床,亟需进一步研究知识获取鲁棒性与安全验证。

原文摘要 · Abstract (English)

The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low resource languages like Persian significantly underserved. To address this gap, this paper introduces Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer grade hardware. As a foundational step toward localized digital healthcare, we first present Gaokerena-V, developed by training a baseline model on a newly curated 90-million-token Persian medical corpus and 20,000 expert-vetted physician Q&A pairs, which improved performance on a translated medical MMLU benchmark from 46.28% to 49.31%. Second, recognizing the critical demands of clinical reasoning, we developed Gaokerena-R by integrating a Chain-of-Thought approach with two novel Reinforcement Learning with AI Feedback (RLAIF) frameworks to optimize preference-based reasoning. Despite utilizing the same baseline architecture and a smaller dataset than Gaokerena-V, Gaokerena-R achieved a superior benchmark score of 52.98%. Furthermore, both models are equipped with custom-developed uncertainty heads that predict the model's confidence in its responses based solely on internal hidden states. While these results demonstrate significant progress in Persian medical language modeling and proactive safety estimation, current performance levels remain insufficient for direct clinical application, highlighting the necessity for further research into robust knowledge acquisition and rigorous safety verification prior to real world deployment.

医学AI小模型波斯语不确定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。