提出可自验证的表征有效性方法,确保智能体在压缩信息下仍能识别决策风险。
Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task Loss
- 通过贝叶斯风险分组定义表征充分性,用总变差阈值实现一次性外部验证
- 将认证转化为任务损失最优停止问题,给出渐近最优的跟踪-停止策略
- 适用于需在低维表征中保障决策可靠性的强化学习场景
在压缩历史表征下行动的智能体面临结构风险:若不同最优动作的历史被表征混淆,则任何基于该表征的规则都存在不可消除的每轮损失,且智能体无法从自身轨迹中察觉。本文构建四层自认证理论:静态层通过贝叶斯风险分组一致性定义决策论充分性,并以精确总变差阈值定价一次性外部验证;序列层将认证建模为任务损失货币下的最优停止问题,通过覆盖线性规划定义环境级认证复杂度常数,证明任意δ-正确策略的任务损失下界,并给出渐近匹配该下界的跟踪-停止策略;边界层提供显式核切换示例,并指出覆盖策略切换或表征修复所需待证开题定理;明确说明固定核保证不适用于表征修订。两个主定理的完整证明见附录。
原文摘要 · Abstract (English)
Agents that act on a compressed representation of their history face a structural risk: if the representation aliases histories with different optimal actions, no rule measurable with respect to the representation can avoid an irreducible per-round loss, and the agent may be unable to detect this from its own transcript. This paper develops a four-layer theory of self-certification of representation adequacy. The static layer defines decision-theoretic adequacy through a Bayes-risk grouping identity and prices a one-shot external verification by an exact total-variation threshold. The sequential layer poses certification as an optimal-stopping problem in the currency of task loss: we define an environment-wise certification complexity constant through a covering linear program, prove an information-task-loss lower bound for every delta-correct strategy, and give a Certification Track-and-Stop policy whose cost matches the bound asymptotically. A final boundary layer gives an explicit kernel-switching example and identifies the open theorem needed to cover policy switching or representation repair; it does not claim that the fixed-kernel guarantees extend to representation revision. The proofs of the two main theorems are given in full in the appendices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。