arXiv:2503.03200cs.CVcs.RO2025-03被引 3

用变压器模型追踪苹果幼果时空位置,准确率达92.4%。

Transformer-Based Spatio-Temporal Association of Apple Fruitlets

  • 基于变压器架构,融合形状与位置特征进行跨时序关联
  • 在商用果园数据上实现92.4%的F1分数,优于所有基线方法
  • 适合小果实追踪场景,尤其适用于多视角、非固定拍摄条件

本文提出一种基于Transformer的方法,用于在不同日期和相机姿态下采集的立体图像中,对苹果幼果进行时空关联。农业领域现有先进关联方法多针对较大作物,依赖高分辨率点云或时间稳定的特征,但这些在田间小型果实上难以获取。为此,我们设计了一种基于Transformer的架构,编码每个幼果的形状与位置信息,并通过交替自注意力与交叉注意力的多层编码器逐步传播与优化特征。实验表明,该方法在商业苹果园数据上实现了92.4%的F1分数,显著优于所有基线方法与消融实验。

原文摘要 · Abstract (English)

In this paper, we present a transformer-based method to spatio-temporally associate apple fruitlets in stereo-images collected on different days and from different camera poses. State-of-the-art association methods in agriculture are dedicated towards matching larger crops using either high-resolution point clouds or temporally stable features, which are both difficult to obtain for smaller fruit in the field. To address these challenges, we propose a transformer-based architecture that encodes the shape and position of each fruitlet, and propagates and refines these features through a series of transformer encoder layers with alternating self and cross-attention. We demonstrate that our method is able to achieve an F1-score of 92.4% on data collected in a commercial apple orchard and outperforms all baselines and ablations.

目标关联苹果追踪视觉定位Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。