用零训练框架让大模型安全准确地标准化医疗术语
An Agentic Model Context Protocol Framework for Medical Concept Standardization
- 基于MCP协议构建零训练映射系统,避免幻觉
- 实现医疗术语实时查词与结构化推理输出
- 适合临床数据标准化团队快速部署使用
观察性医学结果合作计划(OMOP)通用数据模型(CDM)为异构健康数据提供标准化表示,支持大规模多机构研究。利用OMOP CDM进行数据标准化的关键步骤是将源端医疗术语映射到OMOP标准概念,该过程资源消耗大且易出错。尽管大语言模型(LLMs)有潜力辅助此过程,但其易产生幻觉,未经训练和专家验证无法用于临床部署。为此,我们开发了一种基于模型上下文协议(MCP)的零训练、防幻觉映射系统,MCP是一种标准化且安全的框架,使大语言模型能够与外部资源和工具交互。该系统实现了可解释的映射,显著提升效率和准确性,仅需极少投入。它提供实时词汇查询和适用于探索性与生产环境的结构化推理输出。
原文摘要 · Abstract (English)
The Observational Medical Outcomes Partnership (OMOP) common data model (CDM) provides a standardized representation of heterogeneous health data to support large-scale, multi-institutional research. One critical step in data standardization using OMOP CDM is the mapping of source medical terms to OMOP standard concepts, a procedure that is resource-intensive and error-prone. While large language models (LLMs) have the potential to facilitate this process, their tendency toward hallucination makes them unsuitable for clinical deployment without training and expert validation. Here, we developed a zero-training, hallucination-preventive mapping system based on the Model Context Protocol (MCP), a standardized and secure framework allowing LLMs to interact with external resources and tools. The system enables explainable mapping and significantly improves efficiency and accuracy with minimal effort. It provides real-time vocabulary lookups and structured reasoning outputs suitable for immediate use in both exploratory and production environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。