arXiv:2607.09306cs.CLcs.AI2026-07

行为审计需区分暴露与表现,否则评估结果可能颠倒。

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins

  • 区分模型是否被触发(暴露)和行为是否显现(表现)
  • 同一评估下,暴露与表现的AUROC差值达0.2
  • 评估结果受接口输出方式影响,易导致排名反转

行为审计旨在检验语言模型是否按其声称的方式行事,但现有检测分数未区分两个关键目标:模型回复是否在行为诱导条件下生成(暴露),以及该行为是否真实呈现(表现)。在相同720条回复上,对一个1.46亿参数的冻结表示读出审计器与前沿判断器分别基于两类标签进行评估,当目标切换时,两者之间的差距变化约0.2 AUROC。在部署界面下,单一判决结果导致排名反转:审计器在暴露任务上领先(0.804 vs 0.718),但在表现任务上落后(0.690 vs 0.811)。通过匹配输出分辨率——要么让判断器回答特定目标的连续置信度问题,要么对审计器读出进行阈值化——可消除排名反转,但无法消除交互效应,三类分辨率下交互项均显著不为零(0.207, 0.237, 0.169)。目标决定评估工具间的距离,而接口决定该距离是否改变排序。审计器的双曲几何在此无优势。单一行为检测的AUROC报告不完整:只有明确说明估计量、评估者及其输出接口,结果才可比较。

原文摘要 · Abstract (English)

Behavioural auditing asks whether a language model behaves as it claims, but detection scores are reported without separating two targets: whether a reply was produced under a behaviour-inducing condition (exposure) and whether the behaviour surfaced in it (manifestation). Scoring a compact 146-million-parameter auditor's frozen-representation read-out and a frontier judge against each label on the identical 720 replies, the gap between the instruments moves by roughly 0.2 AUROC when the target changes. Under the judge's deployed interface, a single verdict, the ranking reverses: the auditor leads on exposure, 0.804 against 0.718, and trails on manifestation, 0.690 against 0.811. Matching the output resolution from either direction, by asking the judge a target-specific question answered with a continuous confidence score or by thresholding the auditor's read-out, removes the reversal but not the interaction, which excludes zero at all three resolutions (0.207, 0.237 and 0.169). The target governs how far apart the instruments are; the interface governs whether that distance changes their order. The auditor's hyperbolic geometry confers no advantage here. A single behavioural-detection AUROC is under-specified: such claims are comparable only when they state the estimand, the evaluator, and its output interface.

行为审计模型评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。