arXiv:2601.07975cs.CV2026-01

用高效注意力机制解决无人机影像中玉米点定位难题

An Efficient Additive Kolmogorov-Arnold Transformer for Point-Level Maize Localization in Unmanned Aerial Vehicle Imagery

  • 引入柯尔莫哥洛夫-阿诺德网络替代传统模块,增强小目标特征提取能力
  • 在超高清图像上实现62.8% F1分数,计算量降低12.6%,推理速度提升20.7%
  • 适用于高分辨率农业遥感,尤其适合稀疏分布的作物点定位任务

高分辨率无人机摄影测量已成为精准农业关键技术,支持厘米级作物监测与点级植株定位。然而,无人机影像中的点级玉米定位仍面临三大挑战:(1)目标像素比极低,通常不足0.1%;(2)对超过3000×4000像素的超高清图像,传统二次注意力计算成本过高;(3)农田场景特有的稀疏分布与环境变化,通用视觉模型难以应对。为此,本文提出加性柯尔莫哥洛夫-阿诺德变换器(AKT),以帕德柯尔莫哥洛夫-阿诺德网络(PKAN)模块替代标准MLP,提升小目标特征表达能力,并设计PKAN加性注意力(PAA)建模多尺度空间依赖,显著降低计算复杂度。此外,构建了点级玉米定位(PML)数据集,包含1,928张高分辨率无人机图像及约501,000个点标注,覆盖真实田间条件。大量实验表明,AKT平均F1得分达62.8%,优于现有方法4.2%,同时减少12.6%的浮点运算量,推理吞吐量提升20.7%。下游任务中,株距计数均方误差为7.1,植株间距估计均方根误差为1.95–1.97厘米。结果表明,将柯尔莫哥洛夫-阿诺德表示理论与高效注意力机制结合,可有效支撑高分辨率农业遥感应用。

原文摘要 · Abstract (English)

High-resolution UAV photogrammetry has become a key technology for precision agriculture, enabling centimeter-level crop monitoring and point-level plant localization. However, point-level maize localization in UAV imagery remains challenging due to (1) extremely small object-to-pixel ratios, typically less than 0.1%, (2) prohibitive computational costs of quadratic attention on ultra-high-resolution images larger than 3000 x 4000 pixels, and (3) agricultural scene-specific complexities such as sparse object distribution and environmental variability that are poorly handled by general-purpose vision models. To address these challenges, we propose the Additive Kolmogorov-Arnold Transformer (AKT), which replaces conventional multilayer perceptrons with Pade Kolmogorov-Arnold Network (PKAN) modules to enhance functional expressivity for small-object feature extraction, and introduces PKAN Additive Attention (PAA) to model multiscale spatial dependencies with reduced computational complexity. In addition, we present the Point-based Maize Localization (PML) dataset, consisting of 1,928 high-resolution UAV images with approximately 501,000 point annotations collected under real field conditions. Extensive experiments show that AKT achieves an average F1-score of 62.8%, outperforming state-of-the-art methods by 4.2%, while reducing FLOPs by 12.6% and improving inference throughput by 20.7%. For downstream tasks, AKT attains a mean absolute error of 7.1 in stand counting and a root mean square error of 1.95-1.97 cm in interplant spacing estimation. These results demonstrate that integrating Kolmogorov-Arnold representation theory with efficient attention mechanisms offers an effective framework for high-resolution agricultural remote sensing.

玉米定位无人机影像高效注意力农业遥感

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。