arXiv:2604.11680cs.RO2026-04

用扩散模型生成真实感微机器人的显微图像,支持深度依赖的光学效应。

Dual-Control Frequency-Aware Diffusion Model for Depth-Dependent Optical Microrobot Microscopy Image Generation

  • 双控制分支分别处理3D点云和深度网格,实现精确建模。
  • 在仅80张图像/姿态下训练,SSIM提升20.7%。
  • 适合需要高保真数据增强的微机器人系统研发者。

光学镊子驱动的微机器人在细胞操作与微装配中至关重要,但其自主运行依赖精准的三维感知。由于复杂制备工艺和人工标注耗时,大规模高质量显微数据集稀缺。尽管生成式AI可缓解数据不足,现有基于GAN的方法难以复现关键光学特性,尤其是深度相关的衍射与离焦效应。为此,我们提出Du-FreqNet——一种双控制、频域感知的扩散模型,用于物理一致的显微图像合成。该框架包含两个独立的ControlNet分支,分别编码微机器人3D点云与深度特定网格层;引入自适应频域损失,根据距焦平面距离动态调整高低频成分权重。通过可微傅里叶变换监督,模型捕捉了像素空间方法常忽略的物理频谱分布。在有限数据集(如每姿态80张图像)上训练后,模型实现可控的深度依赖图像生成,相比基线SSIM提升20.7%。大量实验表明,Du-FreqNet能有效泛化至未见姿态,显著提升下游任务性能,包括3D位姿与深度估计,从而推动微机器人系统的鲁棒闭环控制。

原文摘要 · Abstract (English)

Optical microrobots actuated by optical tweezers (OT) are important for cell manipulation and microscale assembly, but their autonomous operation depends on accurate 3D perception. Developing such perception systems is challenging because large-scale, high-quality microscopy datasets are scarce, owing to complex fabrication processes and labor-intensive annotation. Although generative AI offers a promising route for data augmentation, existing generative adversarial network (GAN)-based methods struggle to reproduce key optical characteristics, particularly depth-dependent diffraction and defocus effects. To address this limitation, we propose Du-FreqNet, a dual-control, frequency-aware diffusion model for physically consistent microscopy image synthesis. The framework features two independent ControlNet branches to encode microrobot 3D point clouds and depth-specific mesh layers, respectively. We introduce an adaptive frequency-domain loss that dynamically reweights high- and low-frequency components based on the distance to the focal plane. By leveraging differentiable FFT-based supervision, Du-FreqNet captures physically meaningful frequency distributions often missed by pixel-space methods. Trained on a limited dataset (e.g., 80 images per pose), our model achieves controllable, depth-dependent image synthesis, improving SSIM by 20.7% over baselines. Extensive experiments demonstrate that Du-FreqNet generalizes effectively to unseen poses and significantly enhances downstream tasks, including 3D pose and depth estimation, thereby facilitating robust closed-loop control in microrobotic systems.

显微图像生成扩散模型微机器人物理仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。