用大模型直接从噪声图像快速估算蛋白三维构象,提速数十倍。
CryoFastAR: Fast Cryo-EM Ab Initio Reconstruction Made Easy
- 基于多视角特征和渐进式训练,直接从噪点图像预测粒子姿态
- 在真实与合成数据上达到传统方法同等精度,推理速度提升显著
- 适合需要快速构建蛋白初构模型的研究者,尤其擅长低信噪比场景
从无序图像中进行位姿估计是三维重建、机器人学和科学成像的基础。近期的几何基础模型(如DUSt3R)可实现端到端密集三维重建,但在冷冻电镜(cryo-EM)等科学成像领域尚未得到充分探索,尤其是近原子级蛋白重构。在cryo-EM中,从无序粒子图像进行位姿估计与三维重建仍依赖耗时的迭代优化,主要受限于低信噪比(SNR)和对比度传递函数(CTF)引起的失真。本文提出CryoFastAR,首个能直接从cryo-EM噪声图像预测位姿的几何基础模型,实现快速从头重建。通过整合多视图特征,并在大规模模拟的含真实噪声与CTF调制的数据上训练,提升了位姿估计的准确性和泛化能力。为增强训练稳定性,提出渐进式训练策略:先在较简单条件下学习关键特征,再逐步增加难度以提高鲁棒性。实验表明,CryoFastAR在合成与真实数据集上均达到与传统迭代方法相当的质量,同时显著加速推理过程。
原文摘要 · Abstract (English)
Pose estimation from unordered images is fundamental for 3D reconstruction, robotics, and scientific imaging. Recent geometric foundation models, such as DUSt3R, enable end-to-end dense 3D reconstruction but remain underexplored in scientific imaging fields like cryo-electron microscopy (cryo-EM) for near-atomic protein reconstruction. In cryo-EM, pose estimation and 3D reconstruction from unordered particle images still depend on time-consuming iterative optimization, primarily due to challenges such as low signal-to-noise ratios (SNR) and distortions from the contrast transfer function (CTF). We introduce CryoFastAR, the first geometric foundation model that can directly predict poses from Cryo-EM noisy images for Fast ab initio Reconstruction. By integrating multi-view features and training on large-scale simulated cryo-EM data with realistic noise and CTF modulations, CryoFastAR enhances pose estimation accuracy and generalization. To enhance training stability, we propose a progressive training strategy that first allows the model to extract essential features under simpler conditions before gradually increasing difficulty to improve robustness. Experiments show that CryoFastAR achieves comparable quality while significantly accelerating inference over traditional iterative approaches on both synthetic and real datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。