arXiv:2602.00839cs.CV2026-02被引 1

用扩散模型+视觉语义,精准估计透明物体法向。

TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation

  • 基于扩散模型和DINOv3语义融合,提升无纹理透明物表面感知。
  • 在ClearGrasp上误差降24.4%,11.25°精度提升22.8%。
  • 专为实验室透明器皿设计,适合机器人视觉与自动化场景。

单目透明物体法向估计对实验室自动化至关重要,但因复杂的光折射与反射,传统深度和法向传感器常失效。本文提出TransNormal,利用预训练扩散模型进行单步法向回归。为弥补透明表面缺乏纹理的问题,该方法通过交叉注意力机制融合DINOv3的密集视觉语义,提供强几何线索。同时采用多任务学习目标与小波正则化,保持细粒度结构细节。为此任务构建了物理仿真数据集TransNormal-Synthetic,包含高保真法向图。大量实验表明,TransNormal显著优于现有方法:在ClearGrasp上均方误差降低24.4%,11.25°准确率提升22.8%;在ClearPose上均方误差减少15.2%。代码与数据集将公开于https://longxiang-ai.github.io/TransNormal。

原文摘要 · Abstract (English)

Monocular normal estimation for transparent objects is critical for laboratory automation, yet it remains challenging due to complex light refraction and reflection. These optical properties often lead to catastrophic failures in conventional depth and normal sensors, hindering the deployment of embodied AI in scientific environments. We propose TransNormal, a novel framework that adapts pre-trained diffusion priors for single-step normal regression. To handle the lack of texture in transparent surfaces, TransNormal integrates dense visual semantics from DINOv3 via a cross-attention mechanism, providing strong geometric cues. Furthermore, we employ a multi-task learning objective and wavelet-based regularization to ensure the preservation of fine-grained structural details. To support this task, we introduce TransNormal-Synthetic, a physics-based dataset with high-fidelity normal maps for transparent labware. Extensive experiments demonstrate that TransNormal significantly outperforms state-of-the-art methods: on the ClearGrasp benchmark, it reduces mean error by 24.4% and improves 11.25° accuracy by 22.8%; on ClearPose, it achieves a 15.2% reduction in mean error. The code and dataset will be made publicly available at https://longxiang-ai.github.io/TransNormal.

法向估计透明物体扩散模型机器人视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。