arXiv:2501.10774cs.LG2025-01

无标签数据下用特征归因分布监控模型行为,保障AI对齐与性能稳定。

Model Monitoring in the Absence of Labeled Data via Feature Attributions Distributions

  • 基于特征归因分布分析模型行为变化
  • 无需标签即可检测模型偏离或性能下降
  • 适合部署后持续监控的AI系统研发者

模型监控旨在部署后分析人工智能算法并检测其行为变化。本文研究在预测影响现实决策或用户之前进行模型监控,这一阶段的关键特征是测试时缺乏标签数据,导致难以甚至无法计算性能指标。论文围绕两大主题展开:(i) AI对齐,衡量模型行为是否符合人类价值观;(ii) 性能监控,评估模型是否达成特定准确率目标。采用统一方法论,通过分析特征归因分布,利用其理论性质推导出模型监控的保证与洞见,实现无标签条件下的有效监控。

原文摘要 · Abstract (English)

Model monitoring involves analyzing AI algorithms once they have been deployed and detecting changes in their behaviour. This thesis explores machine learning model monitoring ML before the predictions impact real-world decisions or users. This step is characterized by one particular condition: the absence of labelled data at test time, which makes it challenging, even often impossible, to calculate performance metrics. The thesis is structured around two main themes: (i) AI alignment, measuring if AI models behave in a manner consistent with human values and (ii) performance monitoring, measuring if the models achieve specific accuracy goals or desires. The thesis uses a common methodology that unifies all its sections. It explores feature attribution distributions for both monitoring dimensions. Using these feature attribution explanations, we can exploit their theoretical properties to derive and establish certain guarantees and insights into model monitoring.

模型监控特征归因无标签AI对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。