arXiv:2504.15211cs.AIstat.AP2025-04

用概率方法捕捉生成模型的多重风险,让安全评估更全面可信。

Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation

  • 构建解码参数空间的多模态风险分布,量化不同配置下的潜在危害
  • 通过贝叶斯非参数模型发现多个近优操作区域,揭示风险多样性
  • 融合利益相关方偏好,支持动态探索与可解释的安全评估

生成式AI模型的安全评估常将复杂行为简化为单一数值,仅基于一种解码配置,导致尾部风险、群体差异及多个近优操作点被掩盖。本文提出统一框架,通过建模解码参数与提示词空间中危害行为的分布,以尾部聚焦指标量化风险,并整合利益相关方偏好。技术贡献包括:(i) 定义解码拉什蒙集合,即在给定标准下风险近优的参数区域,测量其大小与分歧度;(ii) 设计受利益相关方条件影响的依赖狄利克雷过程混合模型,学习多模态危害表面;(iii) 提出基于贝叶斯深度学习代理的主动采样流程,高效探索参数空间。该方法融合多重性理论、贝叶斯非参数统计与利益相关方对齐的敏感性分析,推动生成模型的可信部署。

原文摘要 · Abstract (English)

Evaluations of generative AI models often collapse nuanced behaviour into a single number computed for a single decoding configuration. Such point estimates obscure tail risks, demographic disparities, and the existence of multiple near-optimal operating points. We propose a unified framework that embraces multiplicity by modelling the distribution of harmful behaviour across the entire space of decoding knobs and prompts, quantifying risk through tail-focused metrics, and integrating stakeholder preferences. Our technical contributions are threefold: (i) we formalise decoding Rashomon sets, regions of knob space whose risk is near-optimal under given criteria and measure their size and disagreement; (ii) we develop a dependent Dirichlet process (DDP) mixture with stakeholder-conditioned stick-breaking weights to learn multi-modal harm surfaces; and (iii) we introduce an active sampling pipeline that uses Bayesian deep learning surrogates to explore knob space efficiently. Our approach bridges multiplicity theory, Bayesian nonparametrics, and stakeholder-aligned sensitivity analysis, paving the way for trustworthy deployment of generative models.

生成模型安全评估贝叶斯方法多模态风险

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。