arXiv:2512.18309cs.LGcs.AI2025-12

让智能体内部自动学习安全约束,避免外部干预。

Embedded Safety-Aligned Intelligence via Differentiable Internal Alignment Embeddings

  • 用可微嵌入直接在智能体内部建模安全对齐机制。
  • 通过反事实推理和注意力传播,降低潜在危害。
  • 适合研究可微对齐与多智能体安全的学者。

我们提出嵌入式安全对齐智能(ESAI),一种用于多智能体强化学习的理论框架,将对齐约束直接嵌入智能体的内部表征中,使用可微的内部对齐嵌入。与外部奖励塑造或事后安全约束不同,内部对齐嵌入是学习到的隐变量,通过反事实推理预测外部伤害,并通过注意力机制和图传播调节策略更新以减少伤害。ESAI整合了四种机制:基于软参考分布计算的可微反事实对齐惩罚、对齐加权感知注意力、支持时序信用分配的海布学习记忆,以及带偏差缓解控制的相似性加权图扩散。我们分析了在利普希茨连续性和谱约束下内部嵌入有界性的稳定条件,讨论了计算复杂度,并研究了收缩行为及公平性-性能权衡等理论性质。本工作定位为可微对齐机制在多智能体系统中的概念性贡献。我们指出了关于收敛性保证、嵌入维度和高维环境扩展等开放理论问题。实证评估留待未来工作。

原文摘要 · Abstract (English)

We introduce Embedded Safety-Aligned Intelligence (ESAI), a theoretical framework for multi-agent reinforcement learning that embeds alignment constraints directly into agents internal representations using differentiable internal alignment embeddings. Unlike external reward shaping or post-hoc safety constraints, internal alignment embeddings are learned latent variables that predict externalized harm through counterfactual reasoning and modulate policy updates toward harm reduction through attention and graph-based propagation. The ESAI framework integrates four mechanisms: differentiable counterfactual alignment penalties computed from soft reference distributions, alignment-weighted perceptual attention, Hebbian associative memory supporting temporal credit assignment, and similarity-weighted graph diffusion with bias mitigation controls. We analyze stability conditions for bounded internal embeddings under Lipschitz continuity and spectral constraints, discuss computational complexity, and examine theoretical properties including contraction behavior and fairness-performance tradeoffs. This work positions ESAI as a conceptual contribution to differentiable alignment mechanisms in multi-agent systems. We identify open theoretical questions regarding convergence guarantees, embedding dimensionality, and extension to high-dimensional environments. Empirical evaluation is left to future work.

多智能体安全对齐可微嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。