用生成数据训练编码器,把随机信号映射到可解释的分子坐标系。
Latent space mapping of interpretable structural coordinates from stochastic single-molecule signals

- 用物理模型生成数据训练对比编码器,学习信号到结构坐标的映射。
- 识别分子只需一次编码,计算量降低千倍以上。
- 对设备差异和构象变化不变,适合跨设备数据整合。
纳米孔是多功能的单分子传感器,但其应用受限于随机穿孔动力学导致的信息失真。本文通过将时域分析转为基于对比编码器的隐空间映射来解决该问题,该编码器仅使用物理信息模型生成的模拟信号进行训练。该编码器将工程化DNA条形码的固态纳米孔信号映射到可解释的分子坐标系统中。学习到的表示对条形码结构参数敏感,但对采集条件和穿孔构象保持不变,支持跨设备数据融合。分子识别仅需一次编码通过,相比对齐方法计算成本降低三个数量级。通过混合物定量、稀有变异检测、共识条形码重建和实时信号获取进行了实验验证。这一从时域分析转向将结构坐标映射到隐空间的转变,重新定义了随机传感器信号的分析范式,使分类直接关联可解释的编码分子信息。
原文摘要 · Abstract (English)
Nanopores are versatile single-molecular sensors, but their utility is fundamentally constrained by stochastic translocation dynamics warping any encoded information. We resolve it by shifting from time-domain analysis to a learned latent-space mapping via a contrastive encoder trained exclusively on simulated signals from a physics-informed model. This encoder maps solid-state nanopore signals of engineered DNA barcodes into an interpretable molecular coordinate system. The learned representation is responsive to structural barcode parameters while remaining invariant to acquisition conditions and translocation conformation, allowing data pooling across devices. Molecule identification requires a single pass through the encoder, reducing computational cost by three orders of magnitude relative to alignment-based methods. We experimentally validate through mixture quantification, rare-variant detection, consensus barcode reconstruction, and real-time signal acquisition. This shift from temporal analysis to mapping structural coordinates into a latent space changes the paradigm behind analyzing stochastic sensor signals by linking classification to interpretable encoded molecular information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。