arXiv:2608.20597cs.SEcs.AI2026-08

军事指挥系统中代理型AI测试面临根本性挑战,现有方法难保实战表现。

Testing and Evaluation of Agentic AI Systems In Military Command and Control

论文配图:Testing and Evaluation of Agentic AI Systems In Military Command and Control
图 1 · 摘自论文原文
  • 识别出支撑测试的八大假设及其四类核心属性
  • 实测结果无法可靠推演至战场实际行为,因代理特性削弱假设
  • 建议通过限定任务范围、运行约束等机制控制风险,适合军方评估人员

代理型AI正被纳入军事指挥与控制(C2)体系,承诺进行严格测试与人工监督。能否兑现承诺取决于其保障论证,需包含可接受性条件、证据及两者间的逻辑关联。通过对240项已记录测试评估实践的结构化分析,覆盖八个评估维度和三个生命周期阶段,识别出八项既有方法对测试对象的假设,分为四类:系统可定义性、稳定性、可组合性和可监管性。代理特性削弱了全部八项假设,影响的是证据与主张之间的推理链条,而非主张或证据本身。因此,测试可能符合流程要求,却无法支持从测试行为到部署行为的推论。本文提出针对前三个假设集群的十项保障主张,并评估现有与新兴方法的应对能力,通过五个C2场景映射操作后果。可监管性未予评估,因其依赖系统稳定性结果及人因测试方法,超出当前范围。现有记录不支持关于系统级行为的广泛结论,但更狭窄的主张仍可成立,前提是具备成熟方法:限定任务边界、基于轨迹的正确性、可执行的运行时约束、可表征的运行间差异。部分证据责任转移至部署阶段,决定是否列装成为持续行为。当无法生成证据时,可通过设定过期条件与责任归属管理残余不确定性。

原文摘要 · Abstract (English)

Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discharged depends on their supporting assurance case, which requires three elements: claims specifying the conditions for acceptability, evidence bearing on those claims, and an argument connecting the two. Through a structured review of 240 documented Testing and Evaluation (T&E) practices, spanning eight evaluation dimensions and three lifecycle stages, we identify eight assumptions that established methods make about their test article, grouped into four clusters: system specifiability, stability, composability, and supervisability. Agentic properties weaken all eight assumptions. This erosion affects the argument connecting evidence to claims, not the claims or evidence themselves. As a result, test results may satisfy process requirements, but they do not warrant the inference from tested to fielded behavior. We derive ten assurance claims for the first three assumption clusters and assess whether current and emerging methods can address each, mapping operational consequences through five C2 scenarios. Supervisability is identified but not assessed here, since evidencing it depends on system stability results and human factors T&E methods beyond the present scope. The documented record does not support broad claims about system-level behavior, but narrower claims remain recoverable in principle, contingent on mature methods: bounded mission envelopes, trajectory-grounded correctness, executable runtime constraints, and characterized run-to-run variance. Part of the evidentiary burden shifts into deployment, making the determination to field a continuing act. Where evidence cannot be generated, the residual uncertainty can be governed through defined expiry conditions and assigned ownership.

AI测试军事AI保障论证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。