arXiv:2512.17956cs.SEcs.AI2025-12

提出多轮置信度校准方法,提升大模型安全与协作平衡

Victor Calibration (VC): Multi-Pass Confidence Calibration and CP4.3 Governance Stress Test under Round-Table Orchestration

  • 通过多轮证据重评估生成递增置信度值T0<T1<T2
  • 在多个Claude模型上实现置信度单调上升且不破坏安全约束
  • 适合关注大模型安全对齐与行为可解释性的研究者

前沿大模型的安全对齐可能导致过度保守,引发回避或虚假拒绝。本文提出轻量级工具包,包含三部分:(1) Victor Calibration (VC),一种通过迭代证据重评估获取递增置信度代理值T(T0<T1<T2)的多轮协议;(2) FD-Lite,一种仅基于行为的现象学审计,使用固定锚定短语和元前缀陷阱以避免拟人化声明;(3) CP4.3,用于测试排名不变性与分配单调性(M6)的治理压力测试。在Claude 4.5系列模型(Haiku、Sonnet no-thinking、Sonnet thinking)及Opus(单次标准UI访问会话)上,观察到置信度轨迹单调上升且未违反安全不变量,CP4.3行为稳定。本研究由单一操作员(n=1)完成,旨在生成假设,明确邀请社区复现、批评与拓展。附有提示模板与成果计划以支持独立验证。

原文摘要 · Abstract (English)

Safety alignment can make frontier LMs overly conservative, degrading collaboration via hedging or false refusals. We present a lightweight toolkit with three parts: (1) Victor Calibration (VC), a multi-pass protocol that elicits a scalar confidence proxy T (T0<T1<T2) through iterative evidence re-evaluation; (2) FD-Lite, a behavior-only phenomenology audit with a fixed anchor phrase and a meta-prefix trap to avoid anthropomorphic claims; and (3) CP4.3, a governance stress test for rank invariance and allocation monotonicity (M6). Across Claude 4.5 models (Haiku, Sonnet no-thinking, Sonnet thinking) and Opus, we observe monotonic VC trajectories without violating safety invariants, and stable CP4.3 behavior. ("Opus" here refers to a single Claude Opus 4.1 session accessed via a standard UI account, as reported in Table 1.) This work was conducted by a single operator (n=1) and is intended as hypothesis-generating; we explicitly invite replication, critique, and extension by the research community. We include prompt templates and an artifact plan to facilitate independent verification.

大模型安全置信度校准治理测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。