arXiv:2411.16702cs.CYcs.LG2024-11被引 1

用临床试验设计方法审计医疗领域语言模型表现

A Clinical Trial Design Approach to Auditing Language Models in Healthcare Setting

  • 将语言模型审计设计为单盲等效试验,对比专家判断
  • 通过严谨样本量计算,最小化数据抽样需求
  • 已在大型公共卫生网络中落地验证,适合医疗AI评估

我们提出一种针对医疗场景部署语言模型的审计机制。该机制借鉴临床试验设计,将语言模型审计视为单盲等效试验,关注点在于模型与领域专家判断的比较。通过该方法,可依据严格的样本量与检验力计算,确保在维持审计完整性和统计可靠性的同时,仅需最少数量的记录抽样。最后,我们在一个大规模公共健康网络的生产环境中提供了实际审计案例。

原文摘要 · Abstract (English)

We present an audit mechanism for language models, with a focus on models deployed in the healthcare setting. Our proposed mechanism takes inspiration from clinical trial design where we posit the language model audit as a single blind equivalence trial, with the comparison of interest being the subject matter experts. We show that using our proposed method, we can follow principled sample size and power calculations, leading to the requirement of sampling minimum number of records while maintaining the audit integrity and statistical soundness. Finally, we provide a real-world example of the audit used in a production environment in a large-scale public health network.

语言模型审计医疗AI临床试验设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。