小到18GB的显微镜数据集,助力药物靶点预测研究
RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy
- 构建压缩版数据集RxRx3-core,保留关键图像与实验条件
- 支持零样本药物靶点预测,覆盖736基因敲除和1674种化合物
- 开源模型嵌入与代码,适合生物医学与机器学习研究者
高内涵筛选(HCS)显微镜数据集已推动对基因与化学扰动下细胞反应的表征,支持基于细胞的药物-靶点相互作用(DTI)推断。然而,表示学习方法在HCS数据上的应用受限于缺乏易用的数据集和可靠基准。为此,我们提出RxRx3-core,一个经精选且压缩的RxRx3数据集子集,以及配套的DTI基准任务。该数据集仅18GB,显著降低大规模HCS数据的使用门槛,同时保留用于评估表示学习模型在零样本DTI预测任务中表现的关键数据。RxRx3-core包含222,601张显微镜图像,涵盖736个CRISPR基因敲除和1,674种化合物在8个浓度下的实验数据。数据集已发布于HuggingFace与Polaris平台,附带预训练嵌入与基准测试代码,确保研究社区的可访问性。通过提供紧凑数据集与稳健基准,我们旨在加速HCS数据表示学习方法的创新,并支持发现新的生物学洞见。
原文摘要 · Abstract (English)
High Content Screening (HCS) microscopy datasets have transformed the ability to profile cellular responses to genetic and chemical perturbations, enabling cell-based inference of drug-target interactions (DTI). However, the adoption of representation learning methods for HCS data has been hindered by the lack of accessible datasets and robust benchmarks. To address this gap, we present RxRx3-core, a curated and compressed subset of the RxRx3 dataset, and an associated DTI benchmarking task. At just 18GB, RxRx3-core significantly reduces the size barrier associated with large-scale HCS datasets while preserving critical data necessary for benchmarking representation learning models against a zero-shot DTI prediction task. RxRx3-core includes 222,601 microscopy images spanning 736 CRISPR knockouts and 1,674 compounds at 8 concentrations. RxRx3-core is available on HuggingFace and Polaris, along with pre-trained embeddings and benchmarking code, ensuring accessibility for the research community. By providing a compact dataset and robust benchmarks, we aim to accelerate innovation in representation learning methods for HCS data and support the discovery of novel biological insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。