arXiv:2607.21343cs.LGcs.AI2026-07

用病理图像和临床数据生成真实基因表达谱,还能解释哪些区域影响了结果。

M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

论文配图:M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
图 1 · 摘自论文原文
  • 用对比学习融合病理图像与临床数据,构建统一潜在表示。
  • 在TCGA数据集上生成的基因表达谱既真实又功能合理。
  • 通过注意力机制可定位影响基因表达的关键病理区域,模型决策可解释。

整合临床元数据、组织病理图像和分子谱等异构生物医学数据,对全面理解疾病至关重要。然而,基因表达数据获取受限于高成本和隐私问题,制约了多模态研究与人工智能应用。我们提出MultiModal Molecular Generation (M$^3$-Gen) 框架,通过将生成对抗网络以病理图像和临床元数据为条件,生成基因表达谱。M$^3$-Gen利用对比学习从临床变量与图像中学习统一潜在表示,并结合两模态嵌入引导生成模型,产出具有生物学一致性的基因表达数据。在TCGA数据集上的评估表明,该方法生成的基因表达谱具有真实性和功能性。尤为重要的是,通过基于注意力机制的多模态融合,M$^3$-Gen具备内在可解释性:能够识别出病理图像中哪些区域最显著影响特定基因表达谱的生成,使模型决策过程可解释。

原文摘要 · Abstract (English)

Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive disease understanding. However, gene expression data acquisition remains constrained by high costs and privacy concerns, limiting its use in multimodal research and AI-driven applications. We present MultiModal Molecular Generation (M$^3$-Gen), a novel framework for the generation of gene expression profiles by conditioning a Generative Adversarial Network on histopathology images and clinical metadata. M$^3$-Gen learns a unified latent representation from the clinical variables and the images, leveraging contrastive learning, and exploits the embeddings of the two modalities to guide a generative model in producing biologically coherent gene expression profiles. Evaluations on the TCGA dataset demonstrate that M$^3$-Gen generates realistic and functionally meaningful gene expression data. Importantly, by integrating multiple modalities in an attention-based mechanism, M$^3$-Gen provides intrinsic explainability: it allows the identification of which regions of the histopathology images most strongly influenced the generation of specific gene expression profiles, making the model's decisions interpretable by design.

多模态生成基因表达可解释性病理图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。