让人类在军事AI测试中全程参与,确保责任可控与系统可信。
Human-centred test and evaluation of military AI
- 将人纳入AI系统整体测试,贯穿研发到部署全周期
- 强调人类角色对系统表现的实际影响,需量化评估
- 适合政策制定者、技术团队和军事操作人员参考
REAIM 2024行动蓝图指出,军事领域中的AI应用应符合伦理且以人为本,人类必须对使用及其后果保持责任与问责。建立严格的测试、验证与评估(TEVV)框架有助于构建强有力的监督机制。在AI系统的开发与部署过程中,需让人类用户全程参与TEVV。传统人因学的人类中心测试方法需针对需持续监控与评估的部署型AI系统进行调整。应将“人类”作为系统组成部分重新定义语言表述,同时需要配套的标准、要求、评估指标与方法。技术与政策制定者之间关于人类中心式TEVV的对话虽长期必要,但必须有明确目标才能产生实效。在整个系统生命周期中推进TEVV至关重要,尤其涉及人类可扩展性及测试规模限制问题。需加强技术与非技术群体间的沟通,使操作员与决策者理解系统使用带来的风险,从而更好指导研发。支持负责任的AI部署的测试评估,必须包含人类对实际作战效能的影响。向使用者与决策者清晰传达TEVV结果,是做出基于风险判断的关键。
原文摘要 · Abstract (English)
The REAIM 2024 Blueprint for Action states that AI applications in the military domain should be ethical and human-centric and that humans must remain responsible and accountable for their use and effects. Developing rigorous test and evaluation, verification and validation (TEVV) frameworks will contribute to robust oversight mechanisms. TEVV in the development and deployment of AI systems needs to involve human users throughout the lifecycle. Traditional human-centred test and evaluation methods from human factors need to be adapted for deployed AI systems that require ongoing monitoring and evaluation. The language around AI-enabled systems should be shifted to inclusion of the human(s) as a component of the system. Standards and requirements supporting this adjusted definition are needed, as are metrics and means to evaluate them. The need for dialogue between technologists and policymakers on human-centred TEVV will be evergreen, but dialogue needs to be initiated with an objective in mind for it to be productive. Development of TEVV throughout system lifecycle is critical to support this evolution including the issue of human scalability and impact on scale of achievable testing. Communication between technical and non technical communities must be improved to ensure operators and policy-makers understand risk assumed by system use and to better inform research and development. Test and evaluation in support of responsible AI deployment must include the effect of the human to reflect operationally realised system performance. Means of communicating the results of TEVV to those using and making decisions regarding the use of AI based systems will be key in informing risk based decisions regarding use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。