arXiv:2504.08812cs.LGcs.CL2025-04被引 4

用模型内部特征识别异常训练信号,提升大模型监督可靠性。

Mechanistic Anomaly Detection for "Quirky" Language Models

  • 通过分析模型内部特征检测训练与测试环境差异。
  • 部分任务下检测器准确率高,但跨模型泛化能力有限。
  • 适合低风险场景,高风险应用需更优检测与评估方法。

随着大语言模型能力提升,对其的监督愈发困难。若模型对监督者未知的因素敏感,可能导致监督失效。本文研究机制性异常检测(MAD),利用模型内部特征识别异常训练信号,以便进一步审查或剔除。我们训练检测器识别测试环境中显著偏离训练环境的数据点,并在多种'古怪'语言模型上测试了不同检测特征与评分规则。结果表明,某些任务下检测器可实现高区分度,但无一检测器能在所有模型和任务上有效。MAD技术在低风险场景中可能可行,但在高风险场景中仍需检测与评估方法的进一步提升。

原文摘要 · Abstract (English)

As LLMs grow in capability, the task of supervising LLMs becomes more challenging. Supervision failures can occur if LLMs are sensitive to factors that supervisors are unaware of. We investigate Mechanistic Anomaly Detection (MAD) as a technique to augment supervision of capable models; we use internal model features to identify anomalous training signals so they can be investigated or discarded. We train detectors to flag points from the test environment that differ substantially from the training environment, and experiment with a large variety of detector features and scoring rules to detect anomalies in a set of ``quirky'' language models. We find that detectors can achieve high discrimination on some tasks, but no detector is effective across all models and tasks. MAD techniques may be effective in low-stakes applications, but advances in both detection and evaluation are likely needed if they are to be used in high stakes settings.

异常检测大模型监督机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。