用物理模型生成显微镜图像,提升微机器人位姿估计的仿真到现实泛化能力。
Physics-Informed Machine Learning for Efficient Sim-to-Real Data Augmentation in Micro-Object Pose Estimation
- 融合波动光学与深度对齐的生成对抗网络,实现高保真图像合成。
- 合成图像SSIM提升35.6%,实时渲染速度0.022秒/帧,位姿估计精度接近真实数据。
- 可泛化至未见姿态,无需额外数据即可支持新微机器人配置的位姿估计。
精确估计光学微机器人位姿对实现高精度目标追踪和自主生物研究至关重要。然而,现有方法严重依赖大规模高质量显微镜图像数据集,而这些数据因微机器人制造复杂性和人工标注耗时,获取困难且成本高昂。数字孪生系统为仿真到现实的数据增强提供了新路径,但现有技术难以复现复杂的光学显微成像现象,如衍射伪影和深度相关的成像特性。本文提出一种新颖的物理信息深度生成学习框架,首次将基于波动光学的物理渲染与深度对齐整合进生成对抗网络(GAN),高效生成用于微机器人位姿估计的高保真显微图像。相比纯人工智能方法,该方法在结构相似性指数(SSIM)上提升35.6%,同时保持实时渲染速度(0.022秒/帧)。基于合成数据训练的卷积神经网络位姿估计算法,在俯仰角和翻滚角上的准确率分别为93.9%和91.9%,仅比仅使用真实数据训练的模型低5.0%和5.4%。此外,该框架可泛化至未见过的姿态,无需额外训练数据即可实现新型微机器人配置的数据增强与鲁棒位姿估计。
原文摘要 · Abstract (English)
Precise pose estimation of optical microrobots is essential for enabling high-precision object tracking and autonomous biological studies. However, current methods rely heavily on large, high-quality microscope image datasets, which are difficult and costly to acquire due to the complexity of microrobot fabrication and the labour-intensive labelling. Digital twin systems offer a promising path for sim-to-real data augmentation, yet existing techniques struggle to replicate complex optical microscopy phenomena, such as diffraction artifacts and depth-dependent imaging.This work proposes a novel physics-informed deep generative learning framework that, for the first time, integrates wave optics-based physical rendering and depth alignment into a generative adversarial network (GAN), to synthesise high-fidelity microscope images for microrobot pose estimation efficiently. Our method improves the structural similarity index (SSIM) by 35.6% compared to purely AI-driven methods, while maintaining real-time rendering speeds (0.022 s/frame).The pose estimator (CNN backbone) trained on our synthetic data achieves 93.9%/91.9% (pitch/roll) accuracy, just 5.0%/5.4% (pitch/roll) below that of an estimator trained exclusively on real data. Furthermore, our framework generalises to unseen poses, enabling data augmentation and robust pose estimation for novel microrobot configurations without additional training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。