arXiv:2505.04927cs.AI2025-05被引 2

用语言片段过滤信念,实现对AI认知状态的精准控制。

Belief Filtering for Epistemic Control in Linguistic State Space

  • 在语义流形框架中,用语言片段动态表示信念状态。
  • 通过内容感知过滤操作,调节智能体在认知转换中的内部状态。
  • 为提升AI安全与对齐提供可解释、可干预的认知治理新路径。

我们研究信念过滤作为一种机制,用于人工智能代理的认知控制,重点在于调控以语言表达形式表示的内部认知状态。该机制基于语义流形框架构建,其中信念状态是自然语言片段的动态、结构化集合。信念过滤器作为对这些片段在各种认知转换中的内容感知操作。本文展示了这种以语言为基础的认知架构所固有的可解释性与模块化特性,直接支持信念过滤,为代理调控提供了原则性方法。研究强调了通过在代理内部语义空间中进行结构化干预,增强AI安全与对齐的潜力,并指出了架构内嵌认知治理的新方向。

原文摘要 · Abstract (English)

We examine belief filtering as a mechanism for the epistemic control of artificial agents, focusing on the regulation of internal cognitive states represented as linguistic expressions. This mechanism is developed within the Semantic Manifold framework, where belief states are dynamic, structured ensembles of natural language fragments. Belief filters act as content-aware operations on these fragments across various cognitive transitions. This paper illustrates how the inherent interpretability and modularity of such a linguistically-grounded cognitive architecture directly enable belief filtering, offering a principled approach to agent regulation. The study highlights the potential for enhancing AI safety and alignment through structured interventions in an agent's internal semantic space and points to new directions for architecturally embedded cognitive governance.

认知控制语义流形AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。