用生态知识生成伪标签,提升光谱树种分类精度。
From Articles to Canopies: Knowledge-Driven Pseudo-Labelling for Tree Species Classification using LLM Experts

- 基于冠层图与生态先验,用LLM生成伪标签。
- 在真实森林数据上提升5.6%分类准确率。
- 适合需要低标注成本的生态遥感研究者。
高光谱树种分类因标签有限且不均衡、光谱混合(多种物种光谱重叠)以及生态异质性(生态系统间差异)而面临挑战。解决这些问题需融合植被的生物与结构特征,如冠层构型和种间互作,而非仅依赖光谱信号。本文提出一种生物启发的半监督深度学习方法,整合多源地球观测数据(高光谱成像与机载激光雷达),结合生态专家知识。该方法基于预构建的冠层图进行生物启发的伪标签生成,实现低训练成本下的高精度分类。同时,利用大语言模型从可靠来源自动提取物种共存先验,并编码为共存矩阵(含物种共现概率)。这些先验被融入伪标签策略,有效引入专家知识。在真实森林数据集上的实验显示,分类精度比最优基线提升5.6%。专家评估表明共存先验准确性高,差异不超过15%。
原文摘要 · Abstract (English)
Hyperspectral tree species classification is challenging due to limited and imbalanced class labels, spectral mixing (overlapping light signatures from multiple species), and ecological heterogeneity (variability among ecological systems). Addressing these challenges requires methods that integrate biological and structural characteristics of vegetation, such as canopy architecture and interspecific interactions, rather than relying solely on spectral signatures. This paper presents a biologically informed, semi-supervised deep learning method that integrates multi-sensor Earth observation data, specifically hyperspectral imaging (HSI) and airborne laser scanning (ALS), with expert, ecological knowledge. The approach relies on biologically inspired pseudo-labelling over a precomputed canopy graph, yielding accurate classification at low training cost. In addition, ecological priors on species cohabitation are automatically derived from reliable sources using large language models (LLMs) and encoded as a cohabitation matrix with likelihoods of species occurring together. These priors are incorporated into the pseudo-labelling strategy, effectively introducing expert knowledge into the model. Experiments on a real-world forest dataset demonstrate 5.6% improvement over the best reference method. Expert evaluation of cohabitation priors reveals high accuracy with differences no larger than 15%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。