评测无人机在城市环境下的多模态决策安全合规性
MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

- 构建融合视觉、语言与动作的综合评估框架
- 17个模型平均合规分仅0.514,严格维度准确率低至0.16
- 揭示视觉与文本输入共同影响决策稳定性
智慧城市空域正将无人飞行器(UAV)从被动感知平台转变为需在观测退化与语言模糊条件下遵循操作规则的网络物理决策者。现有基准侧重感知、导航与推理,但缺乏对关键决策中物理证据、协议约束与动作风险是否耦合的评估。本文提出MulRobBench,一个离线、协议约束的多模态无人机评估基准,集成真实无人机多模态观测、协议级安全策略与动作级网络物理安全。该基准包含3,024个样本,覆盖17个任务类别节点与12个评分维度,分为四个阶段:运行上下文理解、多模态证据仲裁、退化感知推理与风险感知行动规划。评估结合语义评分与结构诊断,涵盖协议合规性、格式合规性、危险动作、解析失败及维度有效性。17个多模态模型中,最佳语义协议决策得分仅0.5141,最严格均值维度准确率为0.1599。受控的20锚点模态消融实验显示,每模型动作选择变化4–15次,证实视觉与文本输入共同影响决策。分析指出模态信任选择、约束提取、强光、数据缺失与操作简写是决策不稳定的主因。MulRobBench为真实运行约束下可信的多模态无人机决策提供了可复现的评测标准。
原文摘要 · Abstract (English)
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。