构建跨物种脑连接组纠错基准,实现人类级精度自动修正
ConnectomeBench2: A Unified Benchmark for Automated Connectomic Proofreading

- 用统一数据集训练视觉变换模型,融合几何与电镜信息
- 在四种物种上达人类水平准确率,性能随数据量提升
- 支持新物种快速扩展,降低标注成本,适合脑科学与AI交叉研究
脑连接组学中,3D神经元重建的分割错误校正(即纠错)是限制全突触分辨率研究的关键瓶颈。本文发布 ConnectomeBench2,一个涵盖超过71.6万条专家标注纠错决策的统一多物种数据集,包含超450万张图像,覆盖小鼠、人类、斑马鱼和果蝇四大开放连接组数据,涵盖分裂与合并两类错误。基于该数据集训练的单个视觉变换模型,采用共享编码器处理网格几何与电子显微图像,在所有物种上对分裂错误校正达到人类水平准确率,并在不同数据规模与模态下表现持续提升。此外,模型在分布内具有良好校准性,分布距离度量可预测未见数据上的性能退化;结合连接组特异性预训练与基于主动学习的样本选择,显著降低新物种与脑区扩展所需的标注工作量。该基准为训练和评估更强大的视觉模型用于连接组纠错提供基础设施。数据与代码已开源:数据在 Hugging Face 上,代码在 GitHub 上。
原文摘要 · Abstract (English)
Proofreading--correcting segmentation errors in 3D brain reconstructions--is the rate-limiting step in synapse-resolution connectomics. We release ConnectomeBench2, a unified multi-species dataset of over 716,485 expert-labeled proofreading decisions with >4,500,000 associated images spanning four major open connectomes (mouse, human, zebrafish, fly), spanning both split and merge error correction. Trained on this dataset, a single Vision Transformer with shared encoders for mesh geometry and electron microscopy reaches human-level accuracy across species for split error correction and merge error identification, with performance scaling with data size and modality. Beyond accuracy, we show that the model is well-calibrated within distribution, that measures of distribution distance predict where calibration and accuracy will degrade on unseen data, and that connectomics-specific pretraining and active learning-based sample selection show potential to substantially reduce the labeling effort needed to extend to new species and brain regions. The benchmark provides the infrastructure to train and evaluate increasingly capable vision models for connectomic proofreading. Data and code availability. The ConnectomeBench2 dataset is released on Hugging Face at https://huggingface.co/datasets/jeffbbrown2/ConnectomeBench2. The accompanying codebase is available on GitHub at https://github.com/timfarkas/ConnectomeBench2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。