利用模拟数据训练蛋白检测模型,提升低温电镜断层成像分析精度。
The microscope is the mask: privileged views and labels from a cryo-ET forward model

- 基于前向模型生成配对视角,增强模型对噪声的不变性。
- 引入蛋白质位置与身份信息,使特征聚焦于蛋白所在区域。
- 无需微调即可在真实数据上超越现有先进模型,适合生物结构分析者。
我们研究了利用模拟数据训练蛋白质注释模型的方法,用于在有限倾角采集且严重受测量算子污染的低温电子断层扫描(cryo-ET)体积中进行分析。首先,利用前向模型带来的畸变生成同一场景的域特定增强配对视图,集成到LeJEPA自监督框架中以实现不变性目标;其次,利用模拟流程提供的蛋白质位置和身份信息,指导模型架构与损失函数设计,使语义信息在密集特征体积中定位至蛋白质位置。所提出的模型CARNIVAL在未进行微调的情况下,于包含多种蛋白类型及两种处理类型的基准数据集上,评估分类与检测任务表现,优于仅使用模拟数据、未采用前向模型配对视图或特权信息的对比模型。
原文摘要 · Abstract (English)
We explore the use of simulated data for training a model for protein annotation in crowded cryo-electron tomography volumes reconstructed from images collected at limited tilt angles and severely corrupted by the measurement operator. Firstly, we leverage the corruptions imposed by the forward model to generate domain-specific augmented paired views of the exact same scene for an invariance objective integrated into the LeJEPA self-supervised training framework. Secondly, we use additional information from the simulation pipeline such as the positions and identity of proteins in the simulated volumes to inform the architecture of the model and the loss function, so that semantic information is localised at protein positions in the resulting dense feature volume. The resulting model, CARNIVAL, is evaluated without finetuning on classification and detection tasks in real tomograms, using a benchmark dataset containing multiple protein types and two tomogram processing types. We show that CARNIVAL outperforms a state-of-the-art model trained using a contrastive objective on simulated data but without forward model-based paired views or privileged information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。