arXiv:2508.10635cs.CV2025-08被引 2

首个融合卫星影像与传感器数据的交互式环境分析模型

ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation

  • 联合分析卫星图像与实时传感器数据,实现多源信息融合推理
  • 在时序与假设推理任务中达BERTF1 0.902,性能媲美或超越领先模型
  • 适合气候研究、城市规划与生态监测领域研究人员使用

从遥感影像理解环境变化对气候韧性、城市规划和生态系统监测至关重要。现有视觉语言模型忽视环境传感器的因果信号,依赖单一来源描述易受风格偏差影响,且缺乏基于场景的交互推理能力。我们提出ChatENV,首个可交互的视觉语言模型,能联合推理卫星图像对与真实世界传感器数据。该框架:(i) 构建包含17.7万张图像的数据库,涵盖197个国家、62类土地利用的15.2万组时序图像对,并附有丰富传感器元数据(如温度、PM10、CO);(ii) 使用GPT4o和Gemini 2.0进行标注,提升风格与语义多样性;(iii) 采用低秩适应(LoRA)微调Qwen-2.5-VL以支持对话功能。ChatENV在时序推理与“如果...会怎样”情景分析中表现优异(如BERTF1达0.902),性能媲美或超越当前最优时序模型,支持交互式场景分析,为基于传感器感知的环境监测提供强大工具。

原文摘要 · Abstract (English)

Understanding environmental changes from remote sensing imagery is vital for climate resilience, urban planning, and ecosystem monitoring. Yet, current vision language models (VLMs) overlook causal signals from environmental sensors, rely on single-source captions prone to stylistic bias, and lack interactive scenario-based reasoning. We present ChatENV, the first interactive VLM that jointly reasons over satellite image pairs and real-world sensor data. Our framework: (i) creates a 177k-image dataset forming 152k temporal pairs across 62 land-use classes in 197 countries with rich sensor metadata (e.g., temperature, PM10, CO); (ii) annotates data using GPT4o and Gemini 2.0 for stylistic and semantic diversity; and (iii) fine-tunes Qwen-2.5-VL using efficient Low-Rank Adaptation (LoRA) adapters for chat purposes. ChatENV achieves strong performance in temporal and "what-if" reasoning (e.g., BERTF1 0.902) and rivals or outperforms state-of-the-art temporal models, while supporting interactive scenario-based analysis. This positions ChatENV as a powerful tool for grounded, sensor-aware environmental monitoring.

环境监测视觉语言模型多模态交互式分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。