用贝叶斯方法自动设计伦理测试,提升自动驾驶系统评估效率与覆盖度。
SEED-SET: Scalable Evolving Experimental Design for System-level Ethical Testing
- 基于分层高斯过程建模客观与主观评价,分离处理不同评估维度。
- 生成的测试用例比基线多2倍,高维空间覆盖率提升25%。
- 适合需兼顾多方价值判断的智能系统伦理评估场景。
随着无人机等自主系统在高风险、以人为本领域部署日益增多,评估其伦理对齐至关重要,否则可能危及人类生命并导致长期决策偏见。当前自动化伦理评测研究不足,主要因缺乏普遍适用的量化指标以及利益相关方主观判断难以解析建模。为此,我们提出SEED-SET,一种融合领域客观评估与利益相关方主观价值判断的贝叶斯实验设计框架。该框架分别使用分层高斯过程建模两类评价,并采用新型采集策略,根据学习到的定性偏好与目标生成符合利益相关方偏好的测试候选。我们在两个应用场景中验证了该方法,结果表明其表现最优。相较于基线,本方法能生成最多两倍的最优测试用例,且在高维搜索空间中的覆盖率提升1.25倍,同时实现探索与利用之间的可解释高效平衡。
原文摘要 · Abstract (English)
As autonomous systems such as drones, become increasingly deployed in high-stakes, human-centric domains, it is critical to evaluate the ethical alignment since failure to do so imposes imminent danger to human lives, and long term bias in decision-making. Automated ethical benchmarking of these systems is understudied due to the lack of ubiquitous, well-defined metrics for evaluation, and stakeholder-specific subjectivity, which cannot be modeled analytically. To address these challenges, we propose SEED-SET, a Bayesian experimental design framework that incorporates domain-specific objective evaluations, and subjective value judgments from stakeholders. SEED-SET models both evaluation types separately with hierarchical Gaussian Processes, and uses a novel acquisition strategy to propose interesting test candidates based on learnt qualitative preferences and objectives that align with the stakeholder preferences. We validate our approach for ethical benchmarking of autonomous agents on two applications and find our method to perform the best. Our method provides an interpretable and efficient trade-off between exploration and exploitation, by generating up to $2\times$ optimal test candidates compared to baselines, with $1.25\times$ improvement in coverage of high dimensional search spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。