让无人机自主理解动态场景,减少人工干预。
AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding
- 基于智能体的任务识别与调度,实现端到端自主决策。
- 零样本下在多种动态场景中达成高质量语义理解。
- 适合需要实时环境感知的无人机应用,如救援与物流。
无人机在物流运输和灾后响应等动态环境中日益重要,但当前任务多依赖人工监控视频并作出操作决策,存在效率低、适应性差的问题。本文提出 AirVista-II——一个面向具身无人机的端到端智能体系统,旨在实现动态场景下的通用语义理解与推理。该系统融合基于智能体的任务识别与调度机制、多模态感知模块以及针对不同时间场景优化的关键帧提取策略,有效捕捉关键场景信息。实验表明,该系统在零样本设置下,于多种基于无人机的动态场景中均实现了高质量的语义理解。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicles (UAVs) are increasingly important in dynamic environments such as logistics transportation and disaster response. However, current tasks often rely on human operators to monitor aerial videos and make operational decisions. This mode of human-machine collaboration suffers from significant limitations in efficiency and adaptability. In this paper, we present AirVista-II -- an end-to-end agentic system for embodied UAVs, designed to enable general-purpose semantic understanding and reasoning in dynamic scenes. The system integrates agent-based task identification and scheduling, multimodal perception mechanisms, and differentiated keyframe extraction strategies tailored for various temporal scenarios, enabling the efficient capture of critical scene information. Experimental results demonstrate that the proposed system achieves high-quality semantic understanding across diverse UAV-based dynamic scenarios under a zero-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。