首个面向微机器人位姿与深度感知的公开数据集,助力显微镜下精准定位。
A Dataset and Benchmarks for Deep Learning-Based Optical Microrobot Pose and Depth Perception
- 构建包含18类微机器人、176种姿态的23万张图像数据集
- ViT在位姿分类中表现最佳,更深模型更适合深度回归
- 数据规模扩大显著提升性能,适合机器人视觉研究者
光学微机器人通过光镊操控,在生物医学中有广泛应用。但由于微机器人透明或对比度低,且工作环境存在噪声和动态变化,其可靠位姿与深度感知仍是核心挑战。公开数据集对实现可复现研究、推动感知模型发展至关重要。标准化评估可保证算法间公平比较。本文提出首个面向光学显微镜下微机器人感知的公开数据集——OpTical MicroRobot (OTMR),包含232,881张图像,覆盖18种微机器人类型与176种不同位姿。我们对八种深度学习模型(包括神经架构搜索生成的结构)在位姿分类与深度回归两个任务上进行了基准测试。结果表明,视觉变换器(ViT)在位姿分类中准确率最高,而深度回归任务更受益于更深网络结构。此外,增大训练数据规模可显著提升两类任务性能,凸显了OTMR作为复杂微尺度环境下鲁棒、泛化感知基础资源的潜力。
原文摘要 · Abstract (English)
Optical microrobots, manipulated via optical tweezers (OT), have broad applications in biomedicine. However, reliable pose and depth perception remain fundamental challenges due to the transparent or low-contrast nature of the microrobots, as well as the noisy and dynamic conditions of the microscale environments in which they operate. An open dataset is crucial for enabling reproducible research, facilitating benchmarking, and accelerating the development of perception models tailored to microscale challenges. Standardised evaluation enables consistent comparison across algorithms, ensuring objective benchmarking and facilitating reproducible research. Here, we introduce the OpTical MicroRobot dataset (OTMR), the first publicly available dataset designed to support microrobot perception under the optical microscope. OTMR contains 232,881 images spanning 18 microrobot types and 176 distinct poses. We benchmarked the performance of eight deep learning models, including architectures derived via neural architecture search (NAS), on two key tasks: pose classification and depth regression. Results indicated that Vision Transformer (ViT) achieve the highest accuracy in pose classification, while depth regression benefits from deeper architectures. Additionally, increasing the size of the training dataset leads to substantial improvements across both tasks, highlighting OTMR's potential as a foundational resource for robust and generalisable microrobot perception in complex microscale environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。