专为脑电报告设计的轻量级语言模型,提升报告理解与生成准确率。
NeuroLex: A Lightweight Domain Language Model for EEG Report Understanding and Generation
- 基于脑电报告文本训练,专注领域语言特征。
- 在摘要生成和术语问答上优于同规模通用模型。
- 适合临床辅助诊断与脑机接口中的语言解码应用。
临床脑电图(EEG)报告包含特定领域的语言规范,通用语言模型难以捕捉。我们提出NeuroLex,一个仅使用哈佛脑电数据库中EEG报告文本训练的轻量级领域自适应语言模型。与现有生物医学语言模型不同,NeuroLex针对脑电报告的语言与诊断特征进行优化,可独立作为文本模型,也可作为多模态脑电-语言系统的解码主干。通过片段损坏预训练及报告润色、段落摘要、术语问答的指令式微调,NeuroLex学习到脑电解释特有的语法与推理模式。全面评估显示,其在相同规模下比通用模型具有更低困惑度、更高提取与摘要准确率、更好标签效率,以及更强的否定句处理与事实幻觉抵抗能力。具备脑电感知的语言主干,NeuroLex连接了生物医学文本建模与脑机接口应用,为可解释、语言驱动的神经解码提供基础。
原文摘要 · Abstract (English)
Clinical electroencephalogram (EEG) reports encode domain-specific linguistic conventions that general-purpose language models (LMs) fail to capture. We introduce NeuroLex, a lightweight domain-adaptive language model trained purely on EEG report text from the Harvard Electroencephalography Database. Unlike existing biomedical LMs, NeuroLex is tailored to the linguistic and diagnostic characteristics of EEG reporting, enabling it to serve as both an independent textual model and a decoder backbone for multimodal EEG-language systems. Using span-corruption pretraining and instruction-style fine-tuning on report polishing, paragraph summarization, and terminology question answering, NeuroLex learns the syntax and reasoning patterns characteristic of EEG interpretation. Comprehensive evaluations show that it achieves lower perplexity, higher extraction and summarization accuracy, better label efficiency, and improved robustness to negation and factual hallucination compared with general models of the same scale. With an EEG-aware linguistic backbone, NeuroLex bridges biomedical text modeling and brain-computer interface applications, offering a foundation for interpretable and language-driven neural decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。