arXiv:2410.19406cs.LG2024-10ICLR被引 8

用生成样本检测大模型行为变化,防微杜渐。

An Auditing Test To Detect Behavioral Shift in Language Models

论文配图:An Auditing Test To Detect Behavioral Shift in Language Models
图 1 · 摘自论文原文
  • 基于假设检验,仅通过模型输出对比判断行为偏移
  • 仅需数百样本即可发现毒性与翻译性能的显著变化
  • 可调灵敏度参数,适合不同场景的持续监控

随着语言模型逼近人类水平性能,对其行为的全面理解变得至关重要,包括能力、偏见、任务表现及与社会价值观的对齐。尽管初期评估(如红队测试和多样化基准)可建立模型行为画像,但后续微调或部署修改可能带来未预期的行为改变。本文提出一种持续性行为偏移审计(BSA)方法。基于近期假设检验研究,该审计测试仅通过模型生成内容即可检测行为变化。测试将基准模型与待测模型的生成结果进行比较,在控制假阳性率的同时提供理论上的变化检测保证。测试包含可配置的容差参数,可根据不同应用场景调节对行为变化的敏感度。我们通过两个案例研究验证该方法:监测(a)毒性水平和(b)翻译性能的变化。结果显示,仅使用数百个样本即可有效检测出行为分布的有意义变化。

原文摘要 · Abstract (English)

As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task performance, and alignment with societal values. Extensive initial evaluations, including red teaming and diverse benchmarking, can establish a model's behavioral profile. However, subsequent fine-tuning or deployment modifications may alter these behaviors in unintended ways. We present a method for continual Behavioral Shift Auditing (BSA) in LMs. Building on recent work in hypothesis testing, our auditing test detects behavioral shifts solely through model generations. Our test compares model generations from a baseline model to those of the model under scrutiny and provides theoretical guarantees for change detection while controlling false positives. The test features a configurable tolerance parameter that adjusts sensitivity to behavioral changes for different use cases. We evaluate our approach using two case studies: monitoring changes in (a) toxicity and (b) translation performance. We find that the test is able to detect meaningful changes in behavior distributions using just hundreds of examples.

行为审计语言模型假设检验持续监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。