Falcon用自然语言指令实现遥感图像14类任务,仅0.7亿参数就达顶尖水平。
Falcon: A Remote Sensing Vision-Language Foundation Model (Technical Report)
- 基于提示的统一框架,支持图像、区域、像素多层级理解
- 在67个数据集上覆盖14项任务,0.7B参数表现超越多数大模型
- 开源7800万高质量遥感指令数据集,适合遥感与多模态研究者
本文提出专用于遥感领域的视觉-语言基础模型Falcon,构建了统一的提示驱动范式,可高效完成复杂遥感任务。Falcon具备图像、区域及像素级理解与推理能力,仅需自然语言指令和遥感图像即可在14类任务中生成文本结果,包括图像分类、目标检测、分割和图像描述等。为支撑Falcon训练并增强其空间与语义表征能力,我们构建了大规模多任务指令微调数据集Falcon_SFT,包含约7800万高质量样本,覆盖560万张多分辨率、多视角遥感图像,配有分层标注并通过人工采样验证确保可靠性。大量对比实验表明,尽管仅有0.7亿参数,Falcon在67个数据集和14项任务上均表现优异。项目代码、数据集及模型权重已开源(https://github.com/TianHuiLab/Falcon),助力开放社区发展。
原文摘要 · Abstract (English)
This paper introduces a holistic vision-language foundation model tailored for remote sensing, named Falcon. Falcon offers a unified, prompt-based paradigm that effectively executes comprehensive and complex remote sensing tasks. Falcon demonstrates powerful understanding and reasoning abilities at the image, region, and pixel levels. Specifically, given simple natural language instructions and remote sensing images, Falcon can produce impressive results in text form across 14 distinct tasks, i.e., image classification, object detection, segmentation, image captioning, and etc. To facilitate Falcon's training and empower its representation capacity to encode rich spatial and semantic information, we developed Falcon_SFT, a large-scale, multi-task, instruction-tuning dataset in the field of remote sensing. The Falcon_SFT dataset consists of approximately 78 million high-quality data samples, covering 5.6 million multi-spatial resolution and multi-view remote sensing images with diverse instructions. It features hierarchical annotations and undergoes manual sampling verification to ensure high data quality and reliability. Extensive comparative experiments are conducted, which verify that Falcon achieves remarkable performance over 67 datasets and 14 tasks, despite having only 0.7B parameters. We release the complete dataset, code, and model weights at https://github.com/TianHuiLab/Falcon, hoping to help further develop the open-source community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。