arXiv:2605.21617cs.LGq-bio.QM2026-05

用Transformer从染色体交互图中精准定位着丝粒位置

$\textit{BlockFormer}$ : Transformer-based inference from interaction maps

论文配图:$\textit{BlockFormer}$ : Transformer-based inference from interaction maps
图 1 · 摘自论文原文
  • 基于Transformer架构处理实体数量和大小不一的交互图
  • 在多种基因组大小的物种中准确恢复着丝粒位置
  • 自研模拟器生成低成本合成数据,提升泛化能力

从交互图(如全基因组染色体构象捕获技术Hi-C)中进行推断可视为一个通用逆问题:给定描述实体间成对相互作用的块状地图,推断一组参数。本文提出一种数据驱动方法,利用地图间的共享结构(如局部模式的全局对齐),同时处理真实数据中实体数量与大小的可变性。方法基于可处理此类变异性得Transformer架构,并设计定制化模拟器生成大量且计算成本低的合成数据用于训练。该方法应用于着丝粒定位任务,在涵盖多种基因组大小的不同物种中均能准确恢复其基因组位置。

原文摘要 · Abstract (English)

Inference from interaction maps, such as centromere identification from genome-wide chromosome conformation capture techniques -- notably Hi-C -- can be formulated as a generic inverse problem: infer a set of parameters given a map summarizing pairwise interactions between entities through blocks of variable numbers and sizes. In this work, we introduce a data-driven approach that leverages shared structure between these maps, such as global alignment between localized patterns, while handling the variability in number and size of entities arising in real-world data. Our approach relies on a transformer architecture capable of handling such variability and a custom simulator to generate abundant, yet computationally cheap synthetic data for training. Applied to the problem of centromere localization, the method accurately recovers their genomic positions across a wide range of species of various genome sizes.

Transformer基因组学逆问题着丝粒定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。