提出决策敏感的暴露控制机制,精准权衡信息揭示与实际收益。
Revelation Control
- 设计决策充分暴露准则,仅揭示能影响关键决策的状态差异。
- 实验证明深层探查在模型中具正向决策价值,且复用训练有更高效率。
- 理论可迁移,但具体参数需针对模型定制,适合对决策机制敏感的研究者。
暴露控制旨在选择价格合理的干预措施,仅在揭示的差异能改变关键决策时才暴露隐藏状态,同时分别计算干预带来的信息价值和实际进展价值。本研究为学习系统构建该理论框架,其中当前信息等价的状态在未来训练中可能产生不同响应并倾向不同动作。框架定义了决策充分暴露与暴露深度,将静态贝叶斯精炼嵌入状态依赖的后续价值,提出精确的成本调整因子分解准则:额外的浅层坐标具有决策非冗余性,当其共享标量摘要的状态位于定价的停止/继续边界两侧时。还提出一种目标无关的模型特定实例化协议,并证明在无限制严重度下,仅有限制停止翻转风险不足以保证正期望效用。在Qwen2.5-7B与Mistral-7B-v0.3上,更深的未来学习探测具有正决策价值;复用训练带来严格等计算量优势。Qwen显示浅层可揭示性存在决策非冗余区间;在Mistral中,仅基于独立开发集拟合的标量继续架构,在独立测试集上仍保持正的族系校正下界,支持标量决策充分性在所测试架构族内的成立。证据表明结构可迁移而非数值可迁移:决策理论、成本会计、后续逻辑与评估协议可转移,而经验代理、系数、阈值甚至所需浅层状态维度可能因系统而异。
原文摘要 · Abstract (English)
Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。