arXiv:2601.21235cs.CLcs.AI2026-01

用风险画像分析大模型社会危害,揭示平均表现下隐藏的极端风险差异。

SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models

  • 将社会危害建模为多维随机变量,分解为偏见、公平性等维度并聚合为累积对数风险
  • 在901个敏感提示上测试11个前沿大模型,发现平均风险相似但尾部暴露差异超2倍
  • 适合关注模型极端风险、推动负责任评估的研究者与监管方

大型语言模型(LLMs)正广泛应用于高风险领域,其罕见但严重的失败可能造成不可逆伤害。现有评估基准常将复杂社会风险简化为均值中心的标量分数,掩盖了分布结构、跨维度交互及最坏情况行为。本文提出社会危害风险画像分析框架(SHARP),实现多维、分布感知的社会危害评估。SHARP将危害建模为多元随机变量,显式分解为偏见、公平性、伦理与认知可靠性,并通过联合失败聚合重构为加性累积对数风险。框架采用风险敏感的分布统计量,以条件风险价值(CVaR95)为主要指标,刻画模型最坏情况行为。在固定语料库(n=901个社会敏感提示)上对11个前沿大模型的应用显示,平均风险相近的模型间尾部暴露与波动性差异超过两倍。各维度尾部行为系统性异质:偏见呈现最强尾部严重性,认知与公平风险居中,伦理错配始终较低;这些模式揭示出模型依赖的异质性失败结构,而标量基准将其混淆。结果表明,负责任的评估与治理需超越标量平均,转向多维、尾部敏感的风险画像。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed in high-stakes domains, where rare but severe failures can result in irreversible harm. However, prevailing evaluation benchmarks often reduce complex social risk to mean-centered scalar scores, thereby obscuring distributional structure, cross-dimensional interactions, and worst-case behavior. This paper introduces Social Harm Analysis via Risk Profiles (SHARP), a framework for multidimensional, distribution-aware evaluation of social harm. SHARP models harm as a multivariate random variable and integrates explicit decomposition into bias, fairness, ethics, and epistemic reliability with a union-of-failures aggregation reparameterized as additive cumulative log-risk. The framework further employs risk-sensitive distributional statistics, with Conditional Value at Risk (CVaR95) as a primary metric, to characterize worst-case model behavior. Application of SHARP to eleven frontier LLMs, evaluated on a fixed corpus of n=901 socially sensitive prompts, reveals that models with similar average risk can exhibit more than twofold differences in tail exposure and volatility. Across models, dimension-wise marginal tail behavior varies systematically across harm dimensions, with bias exhibiting the strongest tail severities, epistemic and fairness risks occupying intermediate regimes, and ethical misalignment consistently lower; together, these patterns reveal heterogeneous, model-dependent failure structures that scalar benchmarks conflate. These findings indicate that responsible evaluation and governance of LLMs require moving beyond scalar averages toward multidimensional, tail-sensitive risk profiling.

大模型评估社会风险尾部风险公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。