arXiv:2601.06223cs.CYcs.AI2026-01被引 7

提出AI代理三支柱模型,保障安全可信的自主运行。

Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness

  • 构建透明、可问责、可信的AI代理框架,分阶段推进自动化。
  • 强调人类监督与渐进验证,避免直接全自动化风险。
  • 适合政策制定者、开发者及关注AI伦理的机构参考。

本文提出一种基于透明性、可问责性和可信性的三支柱模型,用于构建和运营安全可靠的AI代理。该框架融合人机协同、强化学习与协作AI的已有成果,定义了从半自动向全自主演进的路径,主张通过阶段性验证实现安全自治,类比自动驾驶的发展模式。透明性与可问责性被视为建立用户信任、缓解生成式AI常见风险(如幻觉、数据偏见、目标错位)的基础。论文还介绍了三项支撑工作:斯坦福审慎民主实验室推动的公众讨论、安全AI代理联盟的跨行业合作,以及符合三支柱模型的开源代理操作环境工具开发。这些贡献为实现透明运行、价值对齐且维持社会信任的负责任AI代理提供了概念清晰度与实践指导。

原文摘要 · Abstract (English)

This paper presents a conceptual and operational framework for developing and operating safe and trustworthy AI agents based on a Three-Pillar Model grounded in transparency, accountability, and trustworthiness. Building on prior work in Human-in-the-Loop systems, reinforcement learning, and collaborative AI, the framework defines an evolutionary path toward autonomous agents that balances increasing automation with appropriate human oversight. The paper argues that safe agent autonomy must be achieved through progressive validation, analogous to the staged development of autonomous driving, rather than through immediate full automation. Transparency and accountability are identified as foundational requirements for establishing user trust and for mitigating known risks in generative AI systems, including hallucinations, data bias, and goal misalignment, such as the inversion problem. The paper further describes three ongoing work streams supporting this framework: public deliberation on AI agents conducted by the Stanford Deliberative Democracy Lab, cross-industry collaboration through the Safe AI Agent Consortium, and the development of open tooling for an agent operating environment aligned with the Three-Pillar Model. Together, these contributions provide both conceptual clarity and practical guidance for enabling the responsible evolution of AI agents that operate transparently, remain aligned with human values, and sustain societal trust.

AI代理可信AI治理框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。