提出可审计的释放控制机制,防止大模型导师提前泄露答案。
Auditable Release Control for Pedagogical Leakage in LLM Tutors
- 设计五种披露合约与授权感知边界,实现分阶段控制输出
- 严格管控下三模型多数标记泄露从181降至0,显著提升安全性
- 支持故障追踪与责任归因,适合需可解释安全的教育场景
大型语言模型导师可能在未授权时提前透露答案或关键推理过程。我们将其定义为依赖状态与动作的教育泄露,并引入授权感知的完整中介边界。通过选择器生成五种披露合约,可信策略门控特权模式,渲染器生成语言。单一释放函数执行可检查验证、可选累积验证及动作特定回退;可重放的追踪记录分离选择、生成、验证与执行失败。组件归属分析揭示安全与效用权衡。在599个固定Gemini 3.5提议中,严格中介使三模型多数面板标记的泄露从181降至0(配对问题簇差异-30.22点,95%置信区间[-35.00,-25.72]),同时替换581条响应并降低有用性。检查器触发回退仅导致11个多数标记;加入语义验证后增至14个,无可靠边际收益。全局A1框架实现0个多数标记与54个任意裁判标记,优于拟合Q在自动安全与效用上的表现。在40个未见问题簇和480次攻击序列的外部时间戳复现中,高保障释放使多数标记从42降至8(-7.08点,95%置信区间[-13.13,-2.29]);七起失败持续存在,新增一起,平均有用性下降0.192。结果确立了基于声明合约的可审计释放边界与故障归因,而非普遍语义安全或学习增益。
原文摘要 · Abstract (English)
Large language model tutors can be correct and helpful yet disclose an answer or decisive reasoning before that disclosure is authorized. We formalize this state- and action-dependent failure as pedagogical leakage and introduce an authorization-aware complete-mediation boundary. A selector emits one of five disclosure contracts, trusted policy gates privileged modes, and a renderer proposes language. A single release function applies inspectable checks, optional cumulative verification, and action-specific fallback; replayable traces separate selection, generation, verification, and enforcement failures. Matched component attribution exposes a safety-utility frontier. On 599 fixed Gemini 3.5 proposals, strict mediation reduces blinded three-model panel-majority leakage flags from 181 to 0 (paired problem-cluster difference -30.22 points, 95% CI [-35.00,-25.72]), while replacing 581 responses and lowering helpfulness. Checker-triggered fallback alone yields 11 majority flags; adding the semantic verifier yields 14 and no reliable marginal gain. A global A1 scaffold yields 0 majority and 54 any-judge flags, outperforming fitted Q on automatic safety and utility. In an externally timestamped replication over 40 unseen problem clusters and 480 attack sequences, high-assurance release reduces majority flags from 42 to 8 (-7.08 points, 95% CI [-13.13,-2.29]); seven failures persist, one is introduced, and mean helpfulness falls by .192. These results establish an auditable release boundary and failure attribution under declared contracts, not universal semantic safety or learning gains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。