Nyxus是面向大规模图像数据的高效特征提取库,支持多模态生物医学分析。
Nyxus: A Next Generation Image Feature Extraction Library for the Big Data and AI Era
- 从零构建的可扩展特征提取框架,支持2D/3D图像的离线处理
- 覆盖放射组学与细胞分析等多领域,兼容CPU/GPU并行计算
- 提供多种使用方式,适合开发者到无代码用户全链条需求
现代成像设备单次实验可生成数TB至数PB的数据。传统图像分析算法因计算效率不足,常需在鲁棒性与准确性间妥协。深度学习虽显著提升了区域分割精度,但各领域专用特征提取库的分散导致性能对比困难。为此,我们开发了新型特征提取库Nyxus,专为2D/3D图像的大规模、外存处理设计,并经过严格标准验证。其全面的特征集覆盖放射组学与细胞分析等多个生物医学领域,具备跨CPU与GPU的计算可扩展性。Nyxus以多种形态发布:供开发者使用的Python包、命令行工具、无需编码的Napari插件,以及符合OCI规范的容器镜像,适用于云平台与超算环境。此外,它支持程序化调优特征集,实现计算效率与覆盖范围的最优平衡,为新型机器学习与深度学习应用提供方法支持。
原文摘要 · Abstract (English)
Modern imaging instruments can produce terabytes to petabytes of data for a single experiment. The biggest barrier to processing big image datasets has been computational, where image analysis algorithms often lack the efficiency needed to process such large datasets or make tradeoffs in robustness and accuracy. Deep learning algorithms have vastly improved the accuracy of the first step in an analysis workflow (region segmentation), but the expansion of domain specific feature extraction libraries across scientific disciplines has made it difficult to compare the performance and accuracy of extracted features. To address these needs, we developed a novel feature extraction library called Nyxus. Nyxus is designed from the ground up for scalable out-of-core feature extraction for 2D and 3D image data and rigorously tested against established standards. The comprehensive feature set of Nyxus covers multiple biomedical domains including radiomics and cellular analysis, and is designed for computational scalability across CPUs and GPUs. Nyxus has been packaged to be accessible to users of various skill sets and needs: as a Python package for code developers, a command line tool, as a Napari plugin for low to no-code users or users that want to visualize results, and as an Open Container Initiative (OCI) compliant container that can be used in cloud or super-computing workflows aimed at processing large data sets. Further, Nyxus enables a new methodological approach to feature extraction allowing for programmatic tuning of many features sets for optimal computational efficiency or coverage for use in novel machine learning and deep learning applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。