arXiv:2608.02660cs.CYcs.AI2026-08中稿 · AAAI

AI开发者应像信托人一样对用户负有忠实义务,确保长期交互中的安全与信任。

AI Alignment and Fiduciary Obligation

  • 从信托责任出发,提出开发者对用户应尽忠诚、关怀、诚信和坦诚四类义务
  • 识别出用户在长期使用中面临四大风险,并对应到具体责任条款
  • 强调责任独立于实际伤害存在,适用于所有深度交互型AI系统

高级AI助手在建议、决策支持、协作、学习、情感陪伴等多重角色中与用户展开长期互动。当前对齐研究多基于人类关系中的伦理传统(如生物伦理、德性伦理、关怀伦理)。本文聚焦用户-人工智能-开发者三方关系,指出开发者对系统行为、记忆与交互参数具有自由裁量权,因此应承担信托责任。基于商业伦理与法律研究,提出忠诚、关怀、诚信、坦诚四项核心信托义务,可作为对齐标准。这四项义务分别对应用户在长期使用中面临的四大风险,并推导出相应的制度保障措施。该框架将对齐标准建立在开发者对用户的法定责任之上,而非依赖于用户价值的促进,且这些责任不以实际损害为前提,具备普遍适用性。

原文摘要 · Abstract (English)

Advanced AI assistants engage users in extended interactions across a widening range of roles, including advice, decision support, collaboration, learning, emotional support, and companionship among others. Current alignment efforts consider what alignment criteria should govern these relationships, drawing on moral traditions developed for human relationships such as bioethics, virtue ethics, care ethics, and relationship science. This paper considers AI alignment criteria in the user-AI-developer triad, since every user-AI interaction is mediated by a developer who exercises discretionary control over a system's behaviour, memory, and engagement parameters. Drawing on business ethics and legal scholarship, I argue that fiduciary theory applies to extended AI assistant deployment. On this basis, the four canonical fiduciary duties of loyalty, care, good faith, and candour can generate alignment criteria for the developer-user relationship. I map four user-side risks of extended AI assistant deployment to the four duties and specify institutional measures that follow from discharging each duty. The discussion complements existing approaches by grounding alignment criteria in obligations the developer owes the user, rather than in values the user-AI interaction should promote, and by showing that those obligations hold independently of any de facto harm to users.

AI对齐信托责任伦理规范开发者义务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。