评测大模型无人机在6G网络下的安全、鲁棒与效率表现
$α^3$-Bench: A Unified Benchmark of Safety, Robustness, and Efficiency for LLM-Based UAV Agents over 6G Networks
- 构建多轮对话式控制框架,模拟真实6G网络动态变化
- 17个主流大模型在50条任务上测试,仅部分模型保持高成功率
- 提出综合评估指标,涵盖安全、通信成本与资源效率
大型语言模型(LLM)正被用作自主无人机任务的高层控制器。然而,现有评估很少考察此类智能体在下一代网络约束下是否仍能保持安全、协议合规且高效运行。本文提出 $α^3$-Bench,一个用于评估基于LLM的无人机自主性的统一基准,将其视为在动态6G条件下进行多轮对话推理与控制的问题。每个任务被建模为基于LLM的无人机智能体与人类操作员之间的语言中介控制循环,决策需满足严格模式有效性、任务策略、发言交替及安全约束,并适应网络切片、延迟、抖动、丢包率、吞吐量和边缘负载的波动。为反映现代代理工作流,$α^3$-Bench集成了支持工具调用与智能体间协作的双动作层,可评估工具使用一致性与多智能体交互。我们构建了包含113,000个对话式无人机任务的语料库,基于UAVBench场景,使用确定性解码对17个前沿大模型在每场景固定50个任务上进行评估。提出复合 $α^3$ 指标,统一六项核心维度:任务结果、安全策略、工具一致性、交互质量、网络鲁棒性与通信成本,并对每秒和每千令牌进行效率归一化评分。结果表明,尽管部分模型在任务成功与安全合规方面表现良好,但在网络条件劣化时鲁棒性与效率差异显著,凸显了对网络感知与资源高效的基于LLM的无人机智能体的需求。数据集已开源于 GitHub: https://github.com/maferrag/AlphaBench
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used as high level controllers for autonomous Unmanned Aerial Vehicle (UAV) missions. However, existing evaluations rarely assess whether such agents remain safe, protocol compliant, and effective under realistic next generation networking constraints. This paper introduces $α^3$-Bench, a benchmark for evaluating LLM driven UAV autonomy as a multi turn conversational reasoning and control problem operating under dynamic 6G conditions. Each mission is formulated as a language mediated control loop between an LLM based UAV agent and a human operator, where decisions must satisfy strict schema validity, mission policies, speaker alternation, and safety constraints while adapting to fluctuating network slices, latency, jitter, packet loss, throughput, and edge load variations. To reflect modern agentic workflows, $α^3$-Bench integrates a dual action layer supporting both tool calls and agent to agent coordination, enabling evaluation of tool use consistency and multi agent interactions. We construct a large scale corpus of 113k conversational UAV episodes grounded in UAVBench scenarios and evaluate 17 state of the art LLMs using a fixed subset of 50 episodes per scenario under deterministic decoding. We propose a composite $α^3$ metric that unifies six pillars: Task Outcome, Safety Policy, Tool Consistency, Interaction Quality, Network Robustness, and Communication Cost, with efficiency normalized scores per second and per thousand tokens. Results show that while several models achieve high mission success and safety compliance, robustness and efficiency vary significantly under degraded 6G conditions, highlighting the need for network aware and resource efficient LLM based UAV agents. The dataset is publicly available on GitHub : https://github.com/maferrag/AlphaBench
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。