专注疾病理解的嵌入模型,提升相似病症区分能力。
DisEmbed: Transforming Disease Understanding through Embeddings
- 基于疾病专精的合成数据训练,聚焦症状与问答对。
- 在疾病上下文识别与相似病区分上超越现有模型。
- 适合医疗检索增强生成等精准疾病应用。
医学领域广阔多样,现有嵌入模型多面向通用医疗应用,但因泛化范围过广,难以深入理解疾病。为填补此空白,本文提出专攻疾病的嵌入模型DisEmbed。该模型在特制合成数据集上训练,包含疾病描述、症状及疾病相关问答对,特别适配疾病任务。通过疾病专用数据集和三元组评估方法进行基准测试,结果表明,DisEmbed在识别疾病上下文及区分相似疾病方面优于其他模型,尤其在检索增强生成(RAG)任务中表现稳健,具有重要应用价值。
原文摘要 · Abstract (English)
The medical domain is vast and diverse, with many existing embedding models focused on general healthcare applications. However, these models often struggle to capture a deep understanding of diseases due to their broad generalization across the entire medical field. To address this gap, I present DisEmbed, a disease-focused embedding model. DisEmbed is trained on a synthetic dataset specifically curated to include disease descriptions, symptoms, and disease-related Q\&A pairs, making it uniquely suited for disease-related tasks. For evaluation, I benchmarked DisEmbed against existing medical models using disease-specific datasets and the triplet evaluation method. My results demonstrate that DisEmbed outperforms other models, particularly in identifying disease-related contexts and distinguishing between similar diseases. This makes DisEmbed highly valuable for disease-specific use cases, including retrieval-augmented generation (RAG) tasks, where its performance is particularly robust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。