用RGB图像精准估算小麦穗体积,突破了传统方法的精度瓶颈。
Fine-Tuned Vision Transformers Capture Complex Wheat Spike Morphology for Volume Estimation from RGB Images
- 用微调的视觉Transformer模型从2D图像预测三维体积
- 在六视角图像上实现5.08%的误差率和0.97的相关性
- 适合农业表型分析人员快速无损测量作物产量关键特征
从二维RGB图像估计三维形态特征(如小麦穗体积)面临深度信息丢失、投影失真和田间遮挡等挑战。本文利用结构光3D扫描作为真实值,对比多种方法进行非破坏性小麦穗体积估算。小麦穗体积与穗干重高度相关,是果穗效率的关键指标。针对穗部复杂几何结构,比较不同神经网络方法,基准包括2D面积投影法和轴对齐截面重建法。微调的视觉变压器(DINOv2和DINOv3)搭配MLP,在六视角室内图像上分别达到5.08%和4.67%的平均绝对百分比误差(MAPE),相关系数达0.96和0.97,优于微调卷积神经网络(ResNet18/50)、专用骨干网络及两种基线方法。冻结DINO主干时,深度监督的LSTM表现更优;但微调后,高级特征提升使简单MLP超越LSTM。结果表明,物体形状显著影响估测精度,不规则形态如小麦穗对几何方法挑战更大,而深度学习更具鲁棒性。在田间单视角图像上微调DINOv3,MAPE为8.39%,相关系数0.90,提供一种快速、准确、非破坏性的小麦穗体积表型新流程。
原文摘要 · Abstract (English)
Estimating three-dimensional morphological traits such as volume from two-dimensional RGB images presents inherent challenges due to the loss of depth information, projection distortions, and occlusions under field conditions. In this work, we explore multiple approaches for non-destructive volume estimation of wheat spikes using RGB images and structured-light 3D scans as ground truth references. Wheat spike volume is promising for phenotyping as it shows high correlation with spike dry weight, a key component of fruiting efficiency. Accounting for the complex geometry of the spikes, we compare different neural network approaches for volume estimation from 2D images and benchmark them against two conventional baselines: a 2D area-based projection and a geometric reconstruction using axis-aligned cross-sections. Fine-tuned Vision Transformers (DINOv2 and DINOv3) with MLPs achieve the lowest MAPE of 5.08\% and 4.67\% and the highest correlation of 0.96 and 0.97 on six-view indoor images, outperforming fine-tuned CNNs (ResNet18 and ResNet50), wheat-specific backbones, and both baselines. When using frozen DINO backbones, deep-supervised LSTMs outperform MLPs, whereas after fine-tuning, improved high-level representations allow simple MLPs to outperform LSTMs. We demonstrate that object shape significantly impacts volume estimation accuracy, with irregular geometries such as wheat spikes posing greater challenges for geometric methods than for deep learning approaches. Fine-tuning DINOv3 on field-based single side-view images yields a MAPE of 8.39\% and a correlation of 0.90, providing a novel pipeline and a fast, accurate, and non-destructive approach for wheat spike volume phenotyping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。