arXiv:2604.00027cs.CLcs.LG2026-04被引 1

跨国家重症数据无需人工标准化,用翻译实现统一预测。

Multi-lingual Multi-institutional Electronic Health Record based Predictive Model

论文配图:Multi-lingual Multi-institutional Electronic Health Record based Predictive Model
图 1 · 摘自论文原文
  • 用大模型逐词翻译非英文病历为英文,统一语言格式。
  • 在7个国际重症数据库上,跨数据集预测准确率显著提升。
  • 适合做跨国医疗数据建模、少样本迁移学习的研究者。

跨机构大规模电子健康记录(EHR)预测受限于数据模式和编码系统的巨大差异。尽管通用数据模型(CDMs)可标准化记录以支持多机构学习,但手动映射词汇和整合数据成本高且难扩展。基于文本的统一方法通过将原始EHR转换为统一文本形式,可在无需显式标准化的情况下实现联合训练。然而,应用于多国数据时,语言差异成为新障碍。本文研究多语言多机构的EHR预测,目标是在不依赖人工标准化的前提下,实现跨国重症监护数据的联合建模。比较两种策略:(i) 使用多语言编码器直接建模多语言记录,(ii) 通过大语言模型进行词级翻译将非英文记录转为英文。在7个公开重症数据库、10项临床任务及多个预测窗口下,基于翻译的语言对齐表现优于多语言编码器,且多机构模型持续超越需人工特征选择与调和的强基线,也优于单数据集训练。进一步证明,该文本框架结合语言对齐可有效支持少样本微调的迁移学习并带来额外增益。据我们所知,这是首个将多语言多国重症EHR数据集整合为单一预测模型的研究,为无语言壁垒的临床预测及未来全球多机构EHR研究提供了可扩展路径。

原文摘要 · Abstract (English)

Large-scale EHR prediction across institutions is hindered by substantial heterogeneity in schemas and code systems. Although Common Data Models (CDMs) can standardize records for multi-institutional learning, the manual harmonization and vocabulary mapping are costly and difficult to scale. Text-based harmonization provides an alternative by converting raw EHR into a unified textual form, enabling pooled learning without explicit standardization. However, applying this paradigm to multi-national datasets introduces an additional layer of heterogeneity, which is "language" that must be addressed for truly scalable EHRs learning. In this work, we investigate multilingual multi-institutional learning for EHR prediction, aiming to enable pooled training across multinational ICU datasets without manual standardization. We compare two practical strategies for handling language barriers: (i) directly modeling multilingual records with multilingual encoders, and (ii) translating non-English records into English via LLM-based word-level translation. Across seven public ICU datasets, ten clinical tasks with multiple prediction windows, translation-based lingual alignment yields more reliable cross-dataset performance than multilingual encoders. The multi-institutional learning model consistently outperforms strong baselines that require manual feature selection and harmonization, and also surpasses single-dataset training. We further demonstrate that text-based framework with lingual alignment effectively performs transfer learning via few-shot fine-tuning, with additional gains. To our knowledge, this is the first study to aggregate multilingual multinational ICU EHR datasets into one predictive model, providing a scalable path toward language-agnostic clinical prediction and future global multi-institutional EHR research.

多语言电子病历重症预测迁移学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。