用嵌入语言模型解析临床试验嵌入,可生成和操控新试验
ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models
- 基于ELM框架,将大模型对齐临床试验嵌入空间
- 能从嵌入准确描述未见试验,还能生成合理新试验
- 适合医学生成与可解释性研究者使用
文本嵌入已成为各类语言应用的核心。然而,现有方法在解释、探索和反向构建嵌入空间方面仍有限,降低了透明度并限制了生成型应用场景。本文采用近期提出的嵌入语言模型(ELM)方法,将大语言模型对齐临床试验嵌入空间。我们开发了一个开源、领域无关的ELM架构与训练框架,设计了针对临床试验的训练任务,并构建了一个专家验证的合成数据集。通过训练一系列ELM模型,探究任务与训练策略的影响。最终模型ctELM仅凭嵌入即可准确描述和比较未见临床试验,并能从新向量生成合理临床试验。进一步实验表明,生成的试验摘要可响应概念向量上年龄与性别维度的移动。我们的公开ELM实现与实验结果将助力大模型在生物医学等领域的嵌入空间对齐。
原文摘要 · Abstract (English)
Text embeddings have become an essential part of a variety of language applications. However, methods for interpreting, exploring and reversing embedding spaces are limited, reducing transparency and precluding potentially valuable generative use cases. In this work, we align Large Language Models to embeddings of clinical trials using the recently reported Embedding Language Model (ELM) method. We develop an open-source, domain-agnostic ELM architecture and training framework, design training tasks for clinical trials, and introduce an expert-validated synthetic dataset. We then train a series of ELMs exploring the impact of tasks and training regimes. Our final model, ctELM, can accurately describe and compare unseen clinical trials from embeddings alone and produce plausible clinical trials from novel vectors. We further show that generated trial abstracts are responsive to moving embeddings along concept vectors for age and sex of study subjects. Our public ELM implementation and experimental results will aid the alignment of Large Language Models to embedding spaces in the biomedical domain and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。