arXiv:2601.03449cs.RO2026-01被引 1

用物理真实野火数字孪生训练无人机,让其自动追踪火线。

FIRE-VLM: A Vision-Language-Driven Reinforcement Learning Framework for UAV Wildfire Tracking in a Physics-Grounded Fire Digital Twin

  • 在高保真野火数字孪生中,用视觉语言模型引导强化学习。
  • 追踪效率提升6倍,视野内停留时间显著增加。
  • 适合无人系统、灾害监控与复杂环境智能决策研究者。

野火监测需要能在极端视觉退化、快速变化的物理动态和稀缺真实数据条件下自主推理的系统。现有无人机导航方法依赖简化仿真器和监督感知流程,缺乏与物理真实火灾环境互动的具身智能体。我们提出FIRE-VLM,首个在高保真、物理基础的野火数字孪生中端到端训练的视觉-语言模型(VLM)引导强化学习框架。该孪生基于美国地质调查局数字高程模型(DEM)地形、LANDFIRE燃料库存及半物理火势扩散求解器构建,可模拟地形引发的火势突进、风力驱动的加速、烟雾遮挡及动态燃料消耗。在此环境中,采用双视角无人机感知的PPO智能体,由类CLIP的VLM引导。通过单个提示词生成的火灾与烟雾特征语义对齐分数,作为基于势能的奖励塑造信号。贡献包括:(1) 从地理信息系统到仿真的野火数字孪生构建管线;(2) 用于无人机火前线追踪的VLM引导强化学习智能体;(3) 结合物理项与VLM语义的火灾感知奖励设计。在五个数字孪生评估任务中,该VLM引导策略将检测时间缩短至原来的1/6,显著提升火线视野保持时间,据我们所知,是首个在千米尺度、物理真实数字孪生火灾中实现的基于强化学习的无人机野火监测系统。

原文摘要 · Abstract (English)

Wildfire monitoring demands autonomous systems capable of reasoning under extreme visual degradation, rapidly evolving physical dynamics, and scarce real-world training data. Existing UAV navigation approaches rely on simplified simulators and supervised perception pipelines, and lack embodied agents interacting with physically realistic fire environments. We introduce FIRE-VLM, the first end-to-end vision-language model (VLM) guided reinforcement learning (RL) framework trained entirely within a high-fidelity, physics-grounded wildfire digital twin. Built from USGS Digital Elevation Model (DEM) terrain, LANDFIRE fuel inventories, and semi-physical fire-spread solvers, this twin captures terrain-induced runs, wind-driven acceleration, smoke plume occlusion, and dynamic fuel consumption. Within this environment, a PPO agent with dual-view UAV sensing is guided by a CLIP-style VLM. Wildfire-specific semantic alignment scores, derived from a single prompt describing active fire and smoke plumes, are integrated as potential-based reward shaping signals. Our contributions are: (1) a GIS-to-simulation pipeline for constructing wildfire digital twins; (2) a VLM-guided RL agent for UAV firefront tracking; and (3) a wildfire-aware reward design that combines physical terms with VLM semantics. Across five digital-twin evaluation tasks, our VLM-guided policy reduces time-to-detection by up to 6 times, increases time-in-FOV, and is, to our knowledge, the first RL-based UAV wildfire monitoring system demonstrated in kilometer-scale, physics-grounded digital-twin fires.

无人机野火追踪强化学习数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。