首个高光谱地物描述数据集,助力遥感视觉语言模型理解地表信息。
HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models
- 融合光谱数据与像素级文字标注,实现地物语义深度解析
- 在多个基准上使用先进编码器提升分类性能,显著优于传统方法
- 适合遥感、计算机视觉及多模态学习研究者使用
我们提出HyperCap,首个大规模高光谱地物描述数据集,旨在提升遥感应用中模型的性能与有效性。与传统高光谱成像(HSI)基准不同,HyperCap将光谱数据与像素级文本注释结合,促进更深层次的语义理解。该数据集由四个基准数据集构建,并采用自动化与人工相结合的混合标注方式,确保准确性与一致性。基于最先进的编码器和多种融合技术的实证评估显示,分类性能显著提升。结果表明,视觉-语言学习在高光谱成像中具有巨大潜力,使HyperCap成为该领域未来研究的基础性资源。代码与数据集可于 https://github.com/arya-domain/HyperCap 获取。
原文摘要 · Abstract (English)
We introduce HyperCap, the first large-scale hyperspectral captioning dataset designed to enhance model performance and effectiveness in remote sensing applications. Unlike traditional hyperspectral imaging (HSI) benchmarks, HyperCap integrates spectral data with pixel-wise textual annotations, enabling deeper semantic understanding. This dataset enhances model performance in tasks like classification and feature extraction, providing a valuable resource for advanced remote sensing applications. HyperCap is constructed from four benchmark datasets and annotated through a hybrid approach combining automated and manual methods to ensure accuracy and consistency. Empirical evaluations using state-of-the-art encoders and diverse fusion techniques demonstrate significant improvements in classification performance. These results underscore the potential of vision-language learning in HSI and position HyperCap as a foundational dataset for future research in the field. The code and dataset are available at https://github.com/arya-domain/HyperCap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。