arXiv:2510.27163cs.SEcs.AI2025-10被引 1

无需真实标签,通过相对评估判断AI系统风险增减。

MARIA: A Framework for Marginal Risk Assessment without Ground Truth in AI Systems

  • 以相对评估替代绝对风险计算,避免依赖真实标签。
  • 提出可预测性、能力、交互主导性三类相对评估方法。
  • 适合评估长期运行系统改造中的风险与收益,供开发团队参考。

在用AI系统取代现有流程前,需确保改进且不引入额外风险。传统评估依赖两系统的真值,但因结果延迟、成本高昂或数据不全,常不可行,尤其对长期被视为安全的传统系统。更实际的做法并非计算绝对风险,而是比较系统间的差异。为此,我们提出一种边际风险评估框架,摆脱对真值或绝对风险的依赖。该框架强调三类相对评估方法:可预测性、能力与交互主导性。通过从绝对评估转向相对评估,本方法为软件团队提供可操作指导,明确AI在何处提升效果、何处引入新风险,并推动负责任的系统采纳。

原文摘要 · Abstract (English)

Before deploying an AI system to replace an existing process, it must be compared with the incumbent to ensure improvement without added risk. Traditional evaluation relies on ground truth for both systems, but this is often unavailable due to delayed or unknowable outcomes, high costs, or incomplete data, especially for long-standing systems deemed safe by convention. The more practical solution is not to compute absolute risk but the difference between systems. We therefore propose a marginal risk assessment framework, that avoids dependence on ground truth or absolute risk. It emphasizes three kinds of relative evaluation methodology, including predictability, capability and interaction dominance. By shifting focus from absolute to relative evaluation, our approach equips software teams with actionable guidance: identifying where AI enhances outcomes, where it introduces new risks, and how to adopt such systems responsibly.

风险评估相对评价无真值

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。