arXiv:2510.13816q-bio.GNcs.AI2025-10被引 1

构建114万条基因组数据问答与可视化配对数据集,助力生成式AI理解生物数据。

GQVis: A Dataset of Genomics Data Questions and Visualizations for Generative AI

  • 通过整合三类基因组数据库,生成问答与可视化匹配数据
  • 包含114万单问题、62.8万双问题及58.9万链式问题数据
  • 适合训练生物信息学领域生成式AI模型的开发者和研究者

数据可视化是基因组学研究中的基础工具,有助于探索、解释和传播复杂的基因组特征。尽管机器学习模型在将数据转化为有意义的可视化方面展现出潜力,但现有模型缺乏针对特定领域的训练基础。为填补这一空白,我们提出一个框架,生成一组将关于基因组数据的抽象低层问题与对应可视化配对的数据集。该方法基于统计图表的先前工作,适应基因组数据的复杂性及专业表达方式。我们还引入多个关联查询与可视化,并附上设计理由、图注和图像替代文本。使用来自三个不同基因组数据仓库(4DN、ENCODE、Chromoscope)的数据,构建了GQVis数据集,包含114万条单查询数据点、62.8万条查询对和58.9万条查询链。GQVis数据集及生成代码已开源,可访问Hugging Face与GitHub。

原文摘要 · Abstract (English)

Data visualization is a fundamental tool in genomics research, enabling the exploration, interpretation, and communication of complex genomic features. While machine learning models show promise for transforming data into insightful visualizations, current models lack the training foundation for domain-specific tasks. In an effort to provide a foundational resource for genomics-focused model training, we present a framework for generating a dataset that pairs abstract, low-level questions about genomics data with corresponding visualizations. Building on prior work with statistical plots, our approach adapts to the complexity of genomics data and the specialized representations used to depict them. We further incorporate multiple linked queries and visualizations, along with justifications for design choices, figure captions, and image alt-texts for each item in the dataset. We use genomics data retrieved from three distinct genomics data repositories (4DN, ENCODE, Chromoscope) to produce GQVis: a dataset consisting of 1.14 million single-query data points, 628k query pairs, and 589k query chains. The GQVis dataset and generation code are available at https://huggingface.co/datasets/HIDIVE/GQVis and https://github.com/hms-dbmi/GQVis-Generation.

基因组学数据集生成式AI可视化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。