arXiv:2503.20556cs.CL2025-03ACL被引 3

用检索方法解决罗马尼亚语医疗术语匹配难题,提升保险理赔效率

A Retrieval-Based Approach to Medical Procedure Matching in Romanian

  • 基于句子嵌入的检索架构,自动匹配医疗机构与保险公司术语
  • 在罗马尼亚语医疗文本上验证,多语言模型表现优于通用模型
  • 为低资源语言医疗NLP提供可复用方案,适合医保系统优化者

准确将医疗机构使用的医疗程序名称映射到保险公司采用的标准术语,是医疗管理中的关键但复杂任务。命名不一致导致程序分类错误,引发行政低效和保险理赔问题。目前许多公司仍依赖人工映射,存在自动化潜力。本文提出一种基于检索的架构,利用句子嵌入实现罗马尼亚语医疗术语匹配。该任务在低资源语言如罗马尼亚语中尤为困难,因现有预训练模型缺乏医学领域适配。我们评估了罗马尼亚语、多语言及医学专用嵌入模型,识别出最优解决方案。研究成果推动了低资源语言医疗自然语言处理的发展。

原文摘要 · Abstract (English)

Accurately mapping medical procedure names from healthcare providers to standardized terminology used by insurance companies is a crucial yet complex task. Inconsistencies in naming conventions lead to missclasified procedures, causing administrative inefficiencies and insurance claim problems in private healthcare settings. Many companies still use human resources for manual mapping, while there is a clear opportunity for automation. This paper proposes a retrieval-based architecture leveraging sentence embeddings for medical name matching in the Romanian healthcare system. This challenge is significantly more difficult in underrepresented languages such as Romanian, where existing pretrained language models lack domain-specific adaptation to medical text. We evaluate multiple embedding models, including Romanian, multilingual, and medical-domain-specific representations, to identify the most effective solution for this task. Our findings contribute to the broader field of medical NLP for low-resource languages such as Romanian.

医疗NLP术语匹配低资源语言检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。