提出五维可信度评估模型,帮用户正确判断AI系统可靠性。
The Trust Calibration Maturity Model for Characterizing and Communicating Trustworthiness of AI Systems
- 构建包含五个维度的可信度成熟度模型
- 在核科学与地震事件分类中验证了有效性
- 适合评估高风险场景下AI系统的可信度
随着强大AI系统普及,用户需要能力来合理校准对系统的信任。随着系统规模扩大,评估其可信度所需信息越来越难以获取,增加了误用风险。本文提出可信度校准成熟度模型(TCMM),涵盖性能表征、偏差与鲁棒性量化、透明性、安全与可靠性、可用性五个维度,可与系统性能信息一同呈现,帮助用户恰当校准信任,制定要求并追踪进展,识别研究缺口。文中以ChatGPT用于高后果核科学决策和PhaseNet(地震模型集成)用于地震源分类为例进行了演示。
原文摘要 · Abstract (English)
Recent proliferation of powerful AI systems has created a strong need for capabilities that help users to calibrate trust in those systems. As AI systems grow in scale, information required to evaluate their trustworthiness becomes less accessible, presenting a growing risk of using these systems inappropriately. We propose the Trust Calibration Maturity Model (TCMM) to characterize and communicate information about AI system trustworthiness. The TCMM incorporates five dimensions of analytic maturity: Performance Characterization, Bias & Robustness Quantification, Transparency, Safety & Security, and Usability. The TCMM can be presented along with system performance information to (1) help a user to appropriately calibrate trust, (2) establish requirements and track progress, and (3) identify research needs. Here, we discuss the TCMM and demonstrate it on two target tasks: using ChatGPT for high consequence nuclear science determinations, and using PhaseNet (an ensemble of seismic models) for categorizing sources of seismic events.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。