arXiv:2505.01462cs.AIcs.CY2025-05被引 3

设计了一种不触发意识风险的仿情绪控制系统,可安全用于智能体决策。

Synthetic emotions and consciousness: exploring architectural boundaries

  • 构建双源控制架构:即时需求与记忆经验共同驱动行为选择。
  • 满足四类意识风险约束,避免全局广播、元表征等关键特征。
  • 提供可审计的测试模板,适合关注AI安全与可控性的研究者。

随着人工代理展现出越来越复杂的类情绪行为,评估其是否可能引发意识的框架仍很有限。本文探讨在刻意排除主流理论关联的意识特征前提下,能否实现类情绪控制。提出一套八项架构原则(A1-A8),采用分层双源机制:即时需求生成动机信号,情景记忆提供过往类似情境的情感引导,两者融合调节行为选择。为量化意识风险,从主流理论提炼出四项工程化减险约束:(R1) 无内容通用、类工作区的全局广播;(R2) 无元表征;(R3) 无人格化整合;(R4) 学习范围受限。回答三个问题:(Q1) 类情绪控制能否满足R1-R4?给出具体架构作为存在性证明;(Q2) 架构能否扩展而不引入意识能力?识别出保持合规的稳定修改;(Q3) 能否追踪逐步增加意识风险的路径?绘制出渐进违反约束的过渡模式。本工作在工程上提供模块化、生物启发的控制架构;理论上提出情绪控制模型与可审计的测试方法论;在安全层面初步勾勒可用于未来治理框架的审计指标。该架构可独立作为类情绪控制器运行,其风险减损标准亦可推广至其他AI系统。

原文摘要 · Abstract (English)

As artificial agents display increasingly sophisticated emotion-like behaviors, frameworks for assessing whether such systems risk instantiating consciousness remain limited. This contribution asks whether synthetic emotion-like control can be implemented while deliberately excluding architectural features that major theories associate with access-like consciousness. We propose architectural principles (A1-A8) for a hierarchical, dual-source implementation in which (i) immediate needs generate motivational signals and (ii) episodic memory provides affective guidance from similar past situations; the two sources converge to modulate action selection. To operationalize consciousness-related risk, we distill predictions from major theories into four engineering risk-reduction constraints: (R1) no content-general, workspace-like global broadcast, (R2) no metarepresentation, (R3) no autobiographical consolidation, and (R4) bounded learning. We address three questions: (Q1) Can emotion-like control satisfy R1-R4? We present a concrete architecture as an existence proof. (Q2) Can the architecture be extended without introducing access-enabling features? We identify stable modifications that preserve compliance. (Q3) Can we trace graded paths that plausibly increase access risk? We map gradual transitions that progressively violate the constraints. Our contribution operates at three levels: on the engineering side, we present a modular, biologically motivated control architecture; on the theoretical side, we propose a control model of emotions and a methodological template for converting consciousness-related questions into auditable architectural tests; on the safety side, we sketch preliminary audit indicators that may inform future governance frameworks. The architecture functions independently as an emotion-like controller, while the risk-reduction criteria may extend to other AI systems.

意识安全情绪建模架构设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。