arXiv:2505.00810cs.LG2025-05被引 1

用混合检索与神经重排序,自动统一医疗数据中的单位不一致问题。

Scalable Unit Harmonization in Medical Informatics via Bayesian-Optimized Retrieval and Transformer-Based Re-ranking

  • 结合词法与语义检索,再用Transformer模型重排序提升匹配精度。
  • 在75亿条数据上达到MRR 0.9833,rank1精度83.39%,rank5召回94.66%。
  • 适合需要跨机构整合临床数据的研究者和医疗系统开发者。

目的:开发并评估一种可扩展的方法,用于统一大规模临床数据集中的单位不一致问题,解决数据互操作性的关键障碍。材料与方法:设计了一种新型单位统一系统,结合BM25、句子嵌入、贝叶斯优化及基于双向Transformer的二分类器,用于检索与匹配实验室检测条目。系统在Optum Clinformatics Datamart数据集(75亿条记录)上进行评估,采用多阶段流程:过滤、识别、统一建议生成、自动化重排序和人工验证。性能通过平均倒数排名(MRR)及其他标准信息检索指标评估。结果:融合BM25与句子嵌入的混合检索方法(MRR: 0.8833)显著优于仅词法(MRR: 0.7985)或仅嵌入(MRR: 0.5277)的方法。基于Transformer的重排序进一步提升性能(绝对提升0.10),使最终系统MRR达0.9833。系统在rank1时达到83.39%精度,在rank5时召回率达94.66%。讨论:该混合架构有效利用了词法与语义方法的互补优势。重排序模块能修正初始检索因医学术语复杂语义关系导致的错误。结论:本框架为临床数据集中的单位统一提供高效、可扩展的解决方案,减少人工干预,提升准确性。统一后的数据可在不同分析中无缝复用,确保跨医疗系统的一致性,支持更可靠的多中心研究与荟萃分析。

原文摘要 · Abstract (English)

Objective: To develop and evaluate a scalable methodology for harmonizing inconsistent units in large-scale clinical datasets, addressing a key barrier to data interoperability. Materials and Methods: We designed a novel unit harmonization system combining BM25, sentence embeddings, Bayesian optimization, and a bidirectional transformer based binary classifier for retrieving and matching laboratory test entries. The system was evaluated using the Optum Clinformatics Datamart dataset (7.5 billion entries). We implemented a multi-stage pipeline: filtering, identification, harmonization proposal generation, automated re-ranking, and manual validation. Performance was assessed using Mean Reciprocal Rank (MRR) and other standard information retrieval metrics. Results: Our hybrid retrieval approach combining BM25 and sentence embeddings (MRR: 0.8833) significantly outperformed both lexical-only (MRR: 0.7985) and embedding-only (MRR: 0.5277) approaches. The transformer-based reranker further improved performance (absolute MRR improvement: 0.10), bringing the final system MRR to 0.9833. The system achieved 83.39\% precision at rank 1 and 94.66\% recall at rank 5. Discussion: The hybrid architecture effectively leverages the complementary strengths of lexical and semantic approaches. The reranker addresses cases where initial retrieval components make errors due to complex semantic relationships in medical terminology. Conclusion: Our framework provides an efficient, scalable solution for unit harmonization in clinical datasets, reducing manual effort while improving accuracy. Once harmonized, data can be reused seamlessly in different analyses, ensuring consistency across healthcare systems and enabling more reliable multi-institutional studies and meta-analyses.

医疗数据单位统一检索增强Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。