统一评估大模型安全、隐私与鲁棒性的跨维度权衡,揭示现有方法的盲区。
aiXamine: Unified Black-Box Evaluation of Cross-Dimensional Trade-offs in LLM Safety, Security, and Privacy

- 构建黑盒测试平台,集成46项测试覆盖九类服务,自动化评估多维信任度。
- 发现安全强化导致过拒率上升(99.3分对1/3良性请求拒绝),隐私独立于对齐指标。
- 揭示蒸馏引发的鲁棒性崩溃:熵值从56.9骤降至2.6,适合模型安全审计者阅读。
部署中的大语言模型存在跨维度失效:模型安全对齐得分可达99.3,却会拒绝三分之一的良性请求;能力全面提升的同时,隐私分数可能下降21分。现有独立评估安全、安全与隐私的框架无法识别此类现象。我们提出aiXamine,一个统一的黑盒评估平台,将安全、安全与隐私视为相互依赖的属性进行综合评估。该平台通过自动化红队测试流程,在九类服务上执行46项测试,生成从提示级诊断到跨服务权衡分析的分层风险图谱,实现对专有与开源模型在相同条件下的可复现对比。对超过120个LLM开展超5000次测试,完成迄今最大规模联合评估,发现三个单轴评估无法察觉的跨维度现象:第一,安全强化带来可量化的‘安全税’——更强对齐系统性增加过拒率;第二,隐私几乎与其它信任属性正交,未被标准对齐捕获;第三,首次识别并形式化表征‘蒸馏诱导鲁棒性坍塌’:无在线校正的离策略蒸馏会导致熵值从56.9暴跌至2.6,严重破坏同一基础架构的鲁棒性。这些发现结合规模收益递减与类别依赖的安全行为表明,信任度本质是多维的——某一维度进步未必带来其他维度进步,甚至会主动损害,而当前对齐方法仍将其视为单一目标。
原文摘要 · Abstract (English)
The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing one in three benign queries, or improve across every capability metric while losing 21 points in privacy. Existing evaluation frameworks that assess safety, security, and privacy independently cannot detect these patterns. We introduce aiXamine, a unified black-box platform that evaluates LLM trustworthiness across safety, security, and privacy as interdependent properties. aiXamine orchestrates 46 tests across nine services through an automated red-teaming pipeline, producing hierarchical risk profiles, from prompt-level diagnostics to cross-service trade-off analytics, that enable reproducible comparison of proprietary and open-weight systems under identical conditions. Applying aiXamine to over 120 LLMs through more than 5,000 test runs, we conduct the largest joint safety, security, and privacy study to date and uncover three cross-dimensional phenomena invisible to single-axis evaluation. First, safety enforcement incurs a quantifiable safety tax: stronger alignment systematically increases over-refusal, forcing providers to choose between protection and utility. Second, privacy is near-orthogonal to other trustworthiness dimensions and not captured by standard alignment. Third, we identify and formally characterize distillation-induced robustness collapse: off-policy distillation without on-policy correction causes entropy collapse, catastrophically destroying robustness (56.9$\to$2.6) on the same base architecture. These findings, compounded by diminishing returns from scale and category-dependent safety behaviors, demonstrate that trustworthiness is inherently multi-dimensional: progress along one axis does not guarantee, and can actively undermine, progress along others, yet current alignment methods treat it as a single objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。