为神经科定制语言模型,支持本地安全部署。
Building Models of Neurological Language
- 用神经科数据构建专用语料库,支持检索增强生成。
- 实现本地化部署,性能达可复现指标。
- 适合医疗领域研究者与临床辅助系统开发者。
本报告记录了神经科专用语言模型的开发与评估过程。初期聚焦于构建定制模型,后因开源与商业医疗大模型快速进展,转而采用检索增强生成(RAG)与表征模型,实现安全、本地部署。关键贡献包括创建神经科专属数据集(病例报告、问答对、教科书衍生数据)、多词表达提取工具,以及基于图结构的医学术语分析。项目还提供用于本地托管的脚本与Docker容器。报告了性能指标与图社区分析结果,未来工作可探索基于phi-4等开源架构的多模态模型。
原文摘要 · Abstract (English)
This report documents the development and evaluation of domain-specific language models for neurology. Initially focused on building a bespoke model, the project adapted to rapid advances in open-source and commercial medical LLMs, shifting toward leveraging retrieval-augmented generation (RAG) and representational models for secure, local deployment. Key contributions include the creation of neurology-specific datasets (case reports, QA sets, textbook-derived data), tools for multi-word expression extraction, and graph-based analyses of medical terminology. The project also produced scripts and Docker containers for local hosting. Performance metrics and graph community results are reported, with future possible work open for multimodal models using open-source architectures like phi-4.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。