arXiv:2508.14934q-bio.GNcs.LG2025-08被引 1

首个整合拟南芥基因与表型数据的多模态数据集,助力精准解析基因-性状关联。

AGP: A Novel Arabidopsis thaliana Genomics-Phenomics Dataset and its HyperGraph Baseline Benchmarking

  • 构建跨基因表达与表型测量的多模态数据集,实现同一植株的多源信息对齐。
  • 提出超图基准模型,能捕捉基因间高阶相互作用,提升性状预测准确性。
  • 适合植物基因组学、生物信息学及多模态机器学习研究者使用。

解析基因如何控制性状仍是生物学的核心挑战。尽管数据采集技术进步显著,基因到性状的映射能力仍受限。该问题涉及植物育种等多个领域,要求模型能够处理高维、异构且具有生物学结构的数据。然而,当前多数数据集仅包含遗传信息或仅包含表型信息,且表型数据高度异质,许多数据集未能充分捕捉其复杂性。关键问题是这些数据未被整合,无法关联描述同一生物样本,限制了机器学习模型对样本多方面特征的理解,影响相关性学习范围,进而降低预测精度。为此,我们提出拟南芥基因组-表型组(AGP)数据集,这是一个精心整理的多模态数据集,将拟南芥(Arabidopsis thaliana)的基因表达谱与其表型特征测量数据相链接。该数据集支持性状预测和可解释图学习等任务。此外,我们基准测试了传统回归模型与解释性基线,包括一个基于生物学先验知识的超图基线,以验证基因-性状关联。据我们所知,这是首个为同一拟南芥样本提供多模态基因信息与异构性状/表型数据的数据集。通过AGP,我们旨在推动研究社区利用基因信息、基因对间的高阶关系以及多源性状数据,更准确地理解基因型与表型之间的联系。

原文摘要 · Abstract (English)

Understanding which genes control which traits in an organism remains one of the central challenges in biology. Despite significant advances in data collection technology, our ability to map genes to traits is still limited. This genome-to-phenome (G2P) challenge spans several problem domains, including plant breeding, and requires models capable of reasoning over high-dimensional, heterogeneous, and biologically structured data. Currently, however, many datasets solely capture genetic information or solely capture phenotype information. Additionally, phenotype data is very heterogeneous, which many datasets do not fully capture. The critical drawback is that these datasets are not integrated, that is, they do not link with each other to describe the same biological specimens. This limits machine learning models' ability to be informed on the various aspects of these specimens, impacting the breadth of correlations learned, and therefore their ability to make more accurate predictions. To address this gap, we present the Arabidopsis Genomics-Phenomics (AGP) Dataset, a curated multi-modal dataset linking gene expression profiles with phenotypic trait measurements in Arabidopsis thaliana, a model organism in plant biology. AGP supports tasks such as phenotype prediction and interpretable graph learning. In addition, we benchmark conventional regression and explanatory baselines, including a biologically-informed hypergraph baseline, to validate gene-trait associations. To the best of our knowledge, this is the first dataset that provides multi-modal gene information and heterogeneous trait or phenotype data for the same Arabidopsis thaliana specimens. With AGP, we aim to foster the research community towards accurately understanding the connection between genotypes and phenotypes using gene information, higher-order gene pairings, and trait data from several sources.

基因组学表型组多模态数据超图建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。