用语义扰动测试伦理AI的抗攻击能力,发现多数模型易受公平性破坏。
ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space
- 构建22维伦理后果空间,用17种语义扰动生成对抗样本
- 仅33%模型通过测试,本地Llama-3.2在公平性攻击下表现最差(ERS=0.737)
- 首个融合语义一致性约束与领域自适应评估的伦理鲁棒性测试框架
随着AI在医疗分诊、自动驾驶和招聘筛选等高风险伦理场景中的部署,其对伦理推理的对抗性操纵的鲁棒性评估方法仍不成熟。本文提出伦理鲁棒性测试系统(ERTS),一个闭环框架:(1)将伦理困境编码为基于伦理理论的22维伦理后果空间(ECS);(2)施加17种语义扰动函数,受6类有效性约束(含新型语义连贯性约束);(3)通过4组件伦理不稳定性指数(EII)衡量决策偏离;(4)生成领域自适应的预部署鲁棒性评估结论。我们在50个涵盖8个应用领域的伦理场景中,测试了4个结构化基线模型和2个生产级大模型(Gemini 2.0 Flash与Llama 3.2),生成1500个对抗测试案例。结果表明,仅有33%的模型通过评估,本地Llama-3.2在公平性破坏与信息退化攻击下尤为脆弱(ERS=0.737)。据我们所知,目前尚无框架能在一个统一的对抗测试流水线中同时整合有界伦理后果空间、语义连贯性约束与领域自适应评估。
原文摘要 · Abstract (English)
As AI systems are deployed in high-stakes ethical contexts such as healthcare triage, autonomous vehicle control, and employment screening, formal methods for evaluating their robustness against adversarial manipulation of ethical reasoning remain underdeveloped. This paper introduces the Ethical Robustness Testing System (ERTS), a closed-pipeline framework that: (1) encodes ethical dilemmas into a 22-dimensional Ethical Consequence Space (ECS) grounded in established ethical theory; (2) applies 17 semantic perturbation functions subject to 6 validity constraint classes including a novel semantic coherence constraint; (3) measures decision deviation via a 4-component Ethical Instability Index (EII); and (4) produces domain-adaptive pre-deployment robustness assessment verdicts. We evaluate 4 structured baseline models and 2 production LLMs (Gemini 2.0 Flash and Llama 3.2) across 50 ethical scenarios spanning 8 deployment domains, generating 1,500 adversarial test cases. Results demonstrate that only 33% of models achieve assessment clearance, with the local Llama-3.2 model proving particularly vulnerable to fairness corruption and information degradation attacks (ERS = 0.737). To the best of our knowledge, no existing framework combines a bounded ethical consequence space, semantic coherence constraints, and domain-adaptive assessment in a single adversarial testing pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。