arXiv:2603.19822cs.CV2026-03被引 3

评测无人机如何理解简短指令并安全执行复杂飞行任务

HUGE-Bench: A Benchmark for High-Level UAV Vision-Language-Action Tasks

  • 构建基于3D高斯溅射的数字孪生环境,支持真实感渲染与碰撞检测
  • 包含8个高层任务、256万米飞行轨迹,评估过程精度与安全性
  • 适合研究高级无人机自主导航与多阶段行为理解的学者使用

现有无人机视觉语言导航基准主要关注长序列步骤描述与目标导向评估,难以诊断实际操作中对简短高层指令的理解能力。本文提出HUGE-Bench,一个针对高层无人机视觉-语言-动作(HL-VLA)任务的基准,测试智能体能否将简洁语言指令转化为安全的多阶段飞行行为。该基准包含4个真实世界数字孪生场景、8个高层任务和256万米飞行轨迹,基于对齐的3D高斯溅射(3DGS)-Mesh表示,结合逼真渲染与可碰撞几何,支持可扩展生成与碰撞感知评估。引入过程导向与碰撞感知指标,评估过程保真度、终端准确性和安全性。在代表性先进VLA模型上的实验揭示了高层语义补全与安全执行方面的显著差距,表明HUGE-Bench是评估高层无人机自主性的诊断性平台。

原文摘要 · Abstract (English)

Existing UAV vision-language navigation (VLN) benchmarks have enabled language-guided flight, but they largely focus on long, step-wise route descriptions with goal-centric evaluation, making them less diagnostic for real operations where brief, high-level commands must be grounded into safe multi-stage behaviors. We present HUGE-Bench, a benchmark for High-Level UAV Vision-Language-Action (HL-VLA) tasks that tests whether an agent can interpret concise language and execute complex, process-oriented trajectories with safety awareness. HUGE-Bench comprises 4 real-world digital twin scenes, 8 high-level tasks, and 2.56M meters of trajectories, and is built on an aligned 3D Gaussian Splatting (3DGS)-Mesh representation that combines photorealistic rendering with collision-capable geometry for scalable generation and collision-aware evaluation. We introduce process-oriented and collision-aware metrics to assess process fidelity, terminal accuracy, and safety. Experiments on representative state-of-the-art VLA models reveal significant gaps in high-level semantic completion and safe execution, highlighting HUGE-Bench as a diagnostic testbed for high-level UAV autonomy.

无人机视觉语言任务规划数字孪生

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。