一个统一框架让智能体在不懂时能诚实拒绝,并解释缺了什么。
Interpretation, Learning, and Empathy as One Constraint: A Residual-Adequacy Architecture with Accountable Abstention

- 用单一残差量控制理解、决策与共情,不足时自动拒绝
- 能解释人类的‘不知道’类型和机器学习中的认知瓶颈
- 适用于真实与人工智能体,预测可验证
智能体需在当前认知下行动、学习未知内容,并良好建模他人以实现协作。这些能力通常由独立机制实现,但共享一种失效模式:当情境超出当前表征能力时,诚实回应应为有原则的拒绝,说明缺失之处。本文提出一种小型认知架构,其中此类限制源于单一量——内容向量与激活表示框架的残差。解释-决策单元(IDU)通过一组局部表示框架(带私有基)解读内容并决定许可动作;该残差驱动单元行为。残差低且无歧义时发出动作;否则单元重新解读、尝试合理扩展,或以带类型、可追溯的终端状态停止。证明该单元总为终止且确定:对任何输入和固定配置,均在有限步内以唯一终端见证完成,因此拒绝行为自带原因。通过绑定开放参数而不改变机制,同一残差-范围约束在三个尺度重现三种已知现象:‘不知’的分类(带类型的拒绝)、局限于单一共享概念的强制误解(有限共情),以及源自有限聚焦窗口而非预设条件的学习前置依赖(发展性前提)。每种实例均在自然与人工智能体上验证,并提出可证伪预测,表明单一定律可统一建模人机认知局限。该模型贡献了统一性与可问责的拒绝机制,其类型与证据均由构造保证。
原文摘要 · Abstract (English)
An agent must act on the situation before it, learn what it cannot yet represent, and model other agents well enough to coordinate. These faculties are usually realized by separate mechanisms, yet they share a failure mode: the situation can exceed what the agent can currently represent, and the honest response is then a principled refusal that says what was missing. We develop a small cognitive architecture in which these limits arise from a single quantity. An Interpretation-Decision Unit (IDU) interprets a content vector through a family of regimes - local representational frames with private bases - and decides which actions it licenses; a scalar residual of the content against the active regimes' representational scope drives the unit. Low residual with a clean licensing emits an action; otherwise the unit re-interprets, attempts a description-length-justified expansion, or halts with a typed, witnessed terminal. We prove the unit is total and deterministic: for any content and fixed configuration it halts in finitely many bounded-cost steps with a unique terminal witness, so abstention carries its cause by construction. By binding the architecture's open parameters without changing its mechanics, the same residual-against-scope constraint recovers three documented phenomena at three scopes: the typology of not-knowing (typed abstention); a forced misunderstanding between agents, localized to one shared concept and invisible to the agent committing it (bounded empathy); and prerequisite dependence in learning derived from a bounded focus window rather than posited (developmental prerequisites). Each instantiation is worked for a natural and an artificial agent and states a falsifiable prediction, so one constraint can model limits in both human and machine cognition. The account contributes a unification and a notion of accountable abstention, typed and witnessed by construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。