用3D知识蒸馏提升2D图像模型,高效精准估测小麦穗体积。
3D Reconstruction and Knowledge Distillation to Improve Multi-View Image Models to Explore Spike Volume Estimation in Wheat

- 用3D点云网络提取鲁棒几何特征,指导2D图像模型训练。
- 蒸馏后模型误差降低至639.93 mm³,相关性提升至0.82。
- 推理速度从160毫秒降至1.4毫秒,适合田间高通量应用。
准确估计小麦穗体积对产量构成分析和抗逆性评估至关重要,但田间测量仍具挑战。主动式三维传感方法如激光雷达(LiDAR)或飞行时间(ToF)易受植株运动影响,且不适用于户外环境,而三维重建计算开销大。直接使用二维图像处理虽有计算优势,但缺乏显式几何信息。为此,我们提出一种融合2D-3D的混合方法,在训练中引入知识蒸馏,实现仅依赖图像的高效推理。首先,利用基于距离的直方图特征训练一个姿态不变的点云网络,获取鲁棒的几何表示。随后,将3D模型与所提出的多视角图像调节变压器(RT)在集成架构中结合。最后,通过特征或标签蒸馏,将集成模型的知识注入纯图像学生模型。两个蒸馏后的RT将平均绝对误差(MAE)从非蒸馏模型的654.31 mm³降至639.93 mm³和644.62 mm³,相关性从0.76提升至0.77和0.82。同时,推理时间从每穗160毫秒减少至1.4毫秒。蒸馏进一步缓解了体积依赖偏差,并使图像模型的潜在表示更具几何感知能力。结果表明,3D引导的2D Transformer训练可实现可扩展、高效的穗体积估算,适用于高通量田间表型分析。
原文摘要 · Abstract (English)
Accurate estimation of wheat spike volume is important for yield component analysis and stress resilience assessment, yet field-based measurement remains challenging. Active 3D sensing methods such as Light Detection and Ranging (LiDAR) or time-of-flight (ToF) are sensitive to plant motion or poorly suited to outdoor conditions, while 3D reconstructions are computationally expensive. Direct 2D image processing would offer computational advantages, but image-based models lack explicit geometric information. We therefore propose a hybrid 2D-3D approach with knowledge distillation during training while enabling efficient image-only inference. First, we train a rigid-invariant point cloud network using distance-based histogram features to obtain pose-robust geometric representations. We then combine the 3D model with a proposed multi-view image-based regulated Transformer (RT) in an ensemble architecture. Finally, we distill the ensemble knowledge into a purely image-based student model using either feature-based or label-based distillation. The two distilled RTs reduce the mean absolute error (MAE) from 654.31 mm$^3$ of the non-distilled RT to 639.93 mm$^3$ and 644.62 mm$^3$, and increase correlation from 0.76 to 0.77 and 0.82, respectively. At the same time, inference time is reduced from 160 ms to 1.4 ms per spike. Distillation further mitigates volume-dependent bias and reshapes the latent representation of the image model toward a geometry-aware shape. Our results demonstrate that 3D-informed training of a 2D Transformer allows for scalable and efficient spike volume estimation for high-throughput field phenotyping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。