arXiv:2510.04073cs.AI2025-10被引 2

用预测性框架防AI价值观漂移,提升伦理一致性。

Moral Anchor System: A Predictive Framework for AI Value Alignment and Drift Prevention

  • 通过贝叶斯推理+LSTM实时监测并预测AI价值偏移
  • 模拟中可降低80%以上价值漂移事件,误报率仅0.08
  • 适合需高安全性的医疗、金融等关键领域应用

人工智能作为超能力助手已广泛融入各领域,但其行为与人类伦理的对齐问题日益突出。核心风险是价值漂移——因环境变化或学习动态导致AI偏离既定价值观,可能引发效率下降或伦理违规。本文提出道德锚定系统(MAS),一种用于检测、预测和缓解AI代理价值漂移的新框架。该系统结合实时贝叶斯推断监控价值状态、LSTM网络预测漂移趋势,并引入以人类为中心的治理层进行自适应干预。强调低延迟响应(<20毫秒)以预防违规,同时通过人反馈监督微调降低误报和警报疲劳。假设:融合概率漂移检测、预测分析与自适应治理,可在模拟中使价值漂移事件减少80%以上,保持85%检测准确率与0.08的低误报率。在目标错位代理上的严格实验验证了其可扩展性与响应速度。创新点在于其预测与自适应特性,区别于静态对齐方法。贡献包括:(1) MAS架构设计;(2) 强调速度与可用性的实证结果;(3) 跨领域适用性洞察;(4) 开源代码支持复现。

原文摘要 · Abstract (English)

The rise of artificial intelligence (AI) as super-capable assistants has transformed productivity and decision-making across domains. Yet, this integration raises critical concerns about value alignment - ensuring AI behaviors remain consistent with human ethics and intentions. A key risk is value drift, where AI systems deviate from aligned values due to evolving contexts, learning dynamics, or unintended optimizations, potentially leading to inefficiencies or ethical breaches. We propose the Moral Anchor System (MAS), a novel framework to detect, predict, and mitigate value drift in AI agents. MAS combines real-time Bayesian inference for monitoring value states, LSTM networks for forecasting drift, and a human-centric governance layer for adaptive interventions. It emphasizes low-latency responses (<20 ms) to prevent breaches, while reducing false positives and alert fatigue via supervised fine-tuning with human feedback. Our hypothesis: integrating probabilistic drift detection, predictive analytics, and adaptive governance can reduce value drift incidents by 80 percent or more in simulations, maintaining high detection accuracy (85 percent) and low false positive rates (0.08 post-adaptation). Rigorous experiments with goal-misaligned agents validate MAS's scalability and responsiveness. MAS's originality lies in its predictive and adaptive nature, contrasting static alignment methods. Contributions include: (1) MAS architecture for AI integration; (2) empirical results prioritizing speed and usability; (3) cross-domain applicability insights; and (4) open-source code for replication.

AI伦理价值对齐预测系统智能治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。