用大模型辅助脑连接组数据纠错,效果超预期。
ConnectomeBench: Can LLMs Proofread the Connectome?
- 构建多模态基准测试,评估大模型在三类纠错任务中的表现。
- 识别神经元片段类型准确率达52%-82%,远超随机水平。
- 可辅助人工纠错,适合脑科学与AI交叉研究者参考。
连接组学——对生物大脑神经连接的绘制——目前需要大量人力进行数据校对,而这些数据来自成像和机器学习辅助分割。随着利用AI代理自动化重要科研任务的热情高涨,我们探索当前AI系统是否能完成数据校对所需的多项任务。为此,我们提出了ConnectomeBench,一个用于评估大语言模型(LLM)在三个关键校对任务中能力的多模态基准:片段类型识别、分裂错误修正和合并错误检测。基于两个大型开源数据集——小鼠视觉皮层一立方毫米和完整果蝇大脑的专家标注数据,我们评估了包括Claude 3.7/4 Sonnet、o4-mini、GPT-4.1、GPT-4o在内的专有多模态模型,以及InternVL-3、NVLM等开源模型。结果显示,当前模型在片段类型识别任务上表现惊人(平衡准确率52%-82%对比随机水平20%-25%),在二选一或多选分裂错误修正任务中准确率可达75%-85%(对比随机水平50%),但在合并错误识别任务上普遍表现不佳。总体而言,尽管最佳模型仍不及专家水平,但展现出令人期待的潜力,未来或可辅助甚至替代人工校对。项目主页:https://github.com/jffbrwn2/ConnectomeBench,数据集:https://huggingface.co/datasets/jeffbbrown2/ConnectomeBench/tree/main
原文摘要 · Abstract (English)
Connectomics - the mapping of neural connections in an organism's brain - currently requires extraordinary human effort to proofread the data collected from imaging and machine-learning assisted segmentation. With the growing excitement around using AI agents to automate important scientific tasks, we explore whether current AI systems can perform multiple tasks necessary for data proofreading. We introduce ConnectomeBench, a multimodal benchmark evaluating large language model (LLM) capabilities in three critical proofreading tasks: segment type identification, split error correction, and merge error detection. Using expert annotated data from two large open-source datasets - a cubic millimeter of mouse visual cortex and the complete Drosophila brain - we evaluate proprietary multimodal LLMs including Claude 3.7/4 Sonnet, o4-mini, GPT-4.1, GPT-4o, as well as open source models like InternVL-3 and NVLM. Our results demonstrate that current models achieve surprisingly high performance in segment identification (52-82% balanced accuracy vs. 20-25% chance) and binary/multiple choice split error correction (75-85% accuracy vs. 50% chance) while generally struggling on merge error identification tasks. Overall, while the best models still lag behind expert performance, they demonstrate promising capabilities that could eventually enable them to augment and potentially replace human proofreading in connectomics. Project page: https://github.com/jffbrwn2/ConnectomeBench and Dataset https://huggingface.co/datasets/jeffbbrown2/ConnectomeBench/tree/main
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。